Reference ingest text_tracks is not caption coverage: 5 sampled frames
text_tracks come from OCR on deduplicated frame states, at most five frames, not every caption. To know if a Short is captioned, check frames yourself.

text_tracks[] in a Sume reference-ingest manifest is not a full caption transcript. The docs say plainly that these are OCR hits on at most five sampled frames, not caption coverage. If you need to know whether a Short has burned-in captions throughout, do not infer it from the number of text lines. Look at frames across the clip, or work from the spoken words.
Reference ingest is dest first: the docs say it is listed only where SUME_COM_REFERENCE_INGEST_ENABLED permits it, with development auto-on and production opt-in.
What the OCR does
The reference ingest docs describe PP-OCRv5 for Korean and Latin text, run at source resolution on deduplicated frame states and merged across frames into lines. Each line has its text, a normalized box, a span, a confidence, a persistence value, a card_id group and a role_hint. The tool corrects nothing. A line below the confidence threshold is marked needs_verification and the tool returns its native-resolution crop.
| Field | Value | Effect |
|---|---|---|
| ocr.languages | default [ko, en] | scripts read |
| ocr.fps | 0.5 to 2, default 1 | sample rate for text states |
| frames read | at most five | deduplicated states only |
| ocr.min_confidence_attach_crop | 0 to 1, default 0.85 | below it, needs_verification with a crop |
A caption check that works
Ask for a handful of frames at times where speech occurs, then look at them. Video frames takes at[] with 1 to 24 times, returns frames at the source size unless you set max_edge (16 to 2160), and always answers with a 202, so poll. The source must be at most 300 seconds, which covers any Short.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/ref.mp4",
"at": [1, 5, 10, 20, 30, 45],
"max_edge": 960
}Why this matters
Burned-in text can be a hook line, a price, a caption or a platform sticker, and the role_hint is only a hint. If you decide a reference has no captions because three lines came back, you may miss a caption style that appears between the sampled frames. For your own output, you control the text exactly: caption cues give start, end and text without any transcription step.
Limits
The OCR reads Korean and Latin by default and does not read other scripts. It can miss text that is small, low-contrast or animated. The manifest gives you evidence for planning, and the docs advise looking again only at uncertain[] entries, at most once per entry, with a frame at a manifest time.
So the practical rule is to treat text_tracks as hints about on-screen text style: where a hook line sits, how large it is, and whether a price card appears. For a real caption audit, sample frames across the whole Short and compare them against the spoken line at each time.
Sources
Related posts
More in Media tools
- Reframe 16:9 to 9:16: crop fractions for left, center, right
A 16:9 to 9:16 crop is a width of 0.3164 of the frame. Use x 0, 0.3418 or 0.6836 for a left, centered or right subject in Sume video filter.
- Restaurant menu video music: a 30-second brief and cost
What to ask Lyria for under a 30-second restaurant menu video: tempo, instruments, a clean ending and the $0.225 audio cost on Sume with a Timeline render.
- Review reel of six 360p Omni drafts: one timeline render for 10 cents
Join six 360p Omni drafts into one review file with Sume Timeline 1.0 for $0.10, with silence as the spine and contain-fit so nothing is cropped.
- Rounded-corner tile from an AI image with an alpha mask in Pillow
Round the corners of a Sume image with a supersampled ImageDraw mask and putalpha. A radius-to-size table, a smooth edge, and a PNG you can drop on any page.
Written by Sume