Notes on a reference ad: by hand vs one reference-ingest call
What you would log by hand while watching a reference ad, and the manifest Sume's reference ingest returns instead: shots, OCR, audio facts, strip.

A reference ingest turns one clip you already host on media.sume.com (up to 300 seconds) into a timestamped manifest in one call: shots, source-resolution OCR, audio facts, visual boundaries and a labeled overview strip (Reference ingest). Doing the same by hand means pausing and writing each of those down yourself.
The point is not that manual notes are wrong. It is that the manifest is deterministic and shaped for the next step, a brief, a remix or a face swap.
What you write down vs what you get
| Note you would take | Manifest field |
|---|---|
| Where each cut falls, and a frame for it | shots[]: frame-exact cuts, one sharpest keyframe each, palette, luma, motion class |
| On-screen text | text_tracks[]: merged OCR lines with box, span, confidence |
| Is there speech or music, how loud | audio: loudness gate, speech presence, beats; transcript if it transcribes |
| Rough cut candidates | boundaries[]: candidate changes with scores |
| A thumbnail sheet | overview: one strip of up to six labeled tiles |
| Where I am unsure | uncertain[]: the only reasons to look again |
The call
video_url must be a media.sume.com artifact, asset or chat attachment of your workspace, and Idempotency-Key is required. The default mode is sync: a 200 with the manifest, or a 202 job if it is not ready in 30 seconds.
curl -X POST https://api.sume.com/v1/reference-ingest \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ingest-ref-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
"purpose": "brief_format"
}'Availability and billing
Reference ingest is listed only where the feature flag allows it: development auto-on, production opt-in. Billing is the Modal compute the read used times the list rate times 1.25 plus a platform fee, never more than the hold; a transcript adds the STT rate only when speech is detected and transcription is allowed.
What the manifest does not cover
OCR runs on deduplicated states from at most five sampled frames, so text_tracks[] is not caption coverage, and the overview strip is for orientation, not measurement. Anything listed in uncertain[] still needs a second look, which you do with a frame read at the manifest time. A transcript is only billed when the track is not silent and speech is detected; for a brief_format read like the one above, speech.allow_billed_stt defaults to false, so set it to true if you want one.
Sources
Related posts
More in Comparisons
- TTS per 1,000 characters: Sume $0.0475 vs MAI-Voice-2.1 $0.022
Sume TTS lists $0.0475 per 1,000 characters; Microsoft lists MAI-Voice-2.1 at $22 per 1M and Flash at $15 per 1M. List prices side by side, with the job math.
- Veo 3.1 Standard, Fast and Lite per-second prices vs Sume rows
Google lists Veo 3.1 at $0.05 to $0.60 per second. Sume's catalog has no Veo row, so this compares Google's prices with Wan 3.0, Omni Flash and Kling on Sume.
- Waiting for Kling 4.0? Eight Sume video ids by 10-second price
Alternatives you can call today instead of waiting for Kling 4.0: eight Sume video model ids with the price of a 10-second clip, cheapest first.
- Walmart 100MB vs Amazon Sponsored Brands 500MB: bitrate each allows
Walmart caps video at 100MB, Amazon Sponsored Brands at 500MB. The average bitrate each allows for 15 to 90 seconds, plus the Sponsored Brands 1 Mbps floor.
Written by Sume