Reference ingest text_tracks is not caption coverage: 5 sampled frames

text_tracks come from OCR on deduplicated frame states, at most five frames, not every caption. To know if a Short is captioned, check frames yourself.

4 min readSume
All posts

text_tracks[] in a Sume reference-ingest manifest is not a full caption transcript. The docs say plainly that these are OCR hits on at most five sampled frames, not caption coverage. If you need to know whether a Short has burned-in captions throughout, do not infer it from the number of text lines. Look at frames across the clip, or work from the spoken words.

Reference ingest is dest first: the docs say it is listed only where SUME_COM_REFERENCE_INGEST_ENABLED permits it, with development auto-on and production opt-in.

What the OCR does

The reference ingest docs describe PP-OCRv5 for Korean and Latin text, run at source resolution on deduplicated frame states and merged across frames into lines. Each line has its text, a normalized box, a span, a confidence, a persistence value, a card_id group and a role_hint. The tool corrects nothing. A line below the confidence threshold is marked needs_verification and the tool returns its native-resolution crop.

OCR settings in reference ingest (Sume docs read 2026-10-05)
FieldValueEffect
ocr.languagesdefault [ko, en]scripts read
ocr.fps0.5 to 2, default 1sample rate for text states
frames readat most fivededuplicated states only
ocr.min_confidence_attach_crop0 to 1, default 0.85below it, needs_verification with a crop

A caption check that works

Ask for a handful of frames at times where speech occurs, then look at them. Video frames takes at[] with 1 to 24 times, returns frames at the source size unless you set max_edge (16 to 2160), and always answers with a 202, so poll. The source must be at most 300 seconds, which covers any Short.

{
  "video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
  "at": [1, 5, 10, 20, 30, 45],
  "max_edge": 960
}

Why this matters

Burned-in text can be a hook line, a price, a caption or a platform sticker, and the role_hint is only a hint. If you decide a reference has no captions because three lines came back, you may miss a caption style that appears between the sampled frames. For your own output, you control the text exactly: caption cues give start, end and text without any transcription step.

Limits

The OCR reads Korean and Latin by default and does not read other scripts. It can miss text that is small, low-contrast or animated. The manifest gives you evidence for planning, and the docs advise looking again only at uncertain[] entries, at most once per entry, with a frame at a manifest time.

So the practical rule is to treat text_tracks as hints about on-screen text style: where a hook line sits, how large it is, and whether a price card appears. For a real caption audit, sample frames across the whole Short and compare them against the spoken line at each time.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume