Notes on a reference ad: by hand vs one reference-ingest call

What you would log by hand while watching a reference ad, and the manifest Sume's reference ingest returns instead: shots, OCR, audio facts, strip.

3 min readSume
All posts

A reference ingest turns one clip you already host on media.sume.com (up to 300 seconds) into a timestamped manifest in one call: shots, source-resolution OCR, audio facts, visual boundaries and a labeled overview strip (Reference ingest). Doing the same by hand means pausing and writing each of those down yourself.

The point is not that manual notes are wrong. It is that the manifest is deterministic and shaped for the next step, a brief, a remix or a face swap.

What you write down vs what you get

Reference ingest manifest, per the docs
Note you would takeManifest field
Where each cut falls, and a frame for itshots[]: frame-exact cuts, one sharpest keyframe each, palette, luma, motion class
On-screen texttext_tracks[]: merged OCR lines with box, span, confidence
Is there speech or music, how loudaudio: loudness gate, speech presence, beats; transcript if it transcribes
Rough cut candidatesboundaries[]: candidate changes with scores
A thumbnail sheetoverview: one strip of up to six labeled tiles
Where I am unsureuncertain[]: the only reasons to look again

The call

video_url must be a media.sume.com artifact, asset or chat attachment of your workspace, and Idempotency-Key is required. The default mode is sync: a 200 with the manifest, or a 202 job if it is not ready in 30 seconds.

curl -X POST https://api.sume.com/v1/reference-ingest \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ingest-ref-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
    "purpose": "brief_format"
  }'

Availability and billing

Reference ingest is listed only where the feature flag allows it: development auto-on, production opt-in. Billing is the Modal compute the read used times the list rate times 1.25 plus a platform fee, never more than the hold; a transcript adds the STT rate only when speech is detected and transcription is allowed.

What the manifest does not cover

OCR runs on deduplicated states from at most five sampled frames, so text_tracks[] is not caption coverage, and the overview strip is for orientation, not measurement. Anything listed in uncertain[] still needs a second look, which you do with a frame read at the manifest time. A transcript is only billed when the track is not silent and speech is detected; for a brief_format read like the one above, speech.allow_billed_stt defaults to false, so set it to true if you want one.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume