Reference ingest coverage: how many frames OCR read
The reference-ingest coverage block reports frames decoded, frames OCR read and the OCR rate. Use it to decide when on-screen text needs a second look.

The coverage block in a Sume reference-ingest manifest tells you how much of the clip the text reader actually saw: frames_decoded, frames_ocr, ocr_fps, ocr_on_cut_frames and whole_duration. If frames_ocr is small next to the clip length, a title card that flashed between samples may not be in text_tracks[], and you should look at the strip or pull frames before you say the reference has no text.
Episode formats lean on short title cards. YouTube's September 23 post describes Shorts series as seasons and episodes with custom thumbnails (YouTube Blog, read 2026-10-04), so a competitor episode you ingest may carry a one-second card.
What the sampling does
Per the reference ingest docs, ocr.fps is 0.5 to 2 with a default of 1, and OCR runs on deduplicated text states, at most five frames. The docs are explicit that text_tracks[] are OCR hits on at most five sampled frames, not caption coverage. So for a 30-second clip at the default rate there are about 30 candidate instants, and at most five frames are read.
| Field | Type | Use |
|---|---|---|
| frames_decoded | number | How many frames the pass decoded |
| frames_ocr | number | How many frames OCR actually read |
| ocr_fps | number | The sampling rate used, 0.5 to 2 |
| ocr_on_cut_frames | boolean | Whether OCR also ran on cut frames |
| whole_duration | boolean | Whether the read spanned the whole clip |
A three-line verdict
Treat an empty text_tracks[] as unproven when whole_duration is false or frames_ocr is 5 on a text-heavy clip. Raising ocr.fps toward 2 samples more instants but does not lift the five-frame ceiling the docs describe, so read the strip or extract frames at the times you care about.
coverage = {
"frames_decoded": 61,
"frames_ocr": 5,
"ocr_fps": 1,
"ocr_on_cut_frames": True,
"whole_duration": True,
}
text_tracks = []
if not coverage["whole_duration"]:
verdict = "partial read: do not trust an empty text list"
elif coverage["frames_ocr"] >= 5 and not text_tracks:
verdict = "five frames read, none had text: check the strip"
elif not text_tracks:
verdict = "whole clip read, no text found"
else:
verdict = f"{len(text_tracks)} text tracks"
print(verdict)
Reference ingest is dest first and production opt-in. When a line is read at low confidence it comes back as needs_verification with a native crop; see OCR crops. Extra frames come from video frames.
Sources
Related posts
More in Media tools
- Reference ingest purpose: qa or remix decides who transcribes
Reference ingest purpose defaults to reference_remix, which transcribes speech. brief_format, face_swap and qa do not. An explicit allow_billed_stt wins.
- Reference ingest shot record: cut types, palette, luma, motion
Each reference-ingest shot carries cut_out type (hard, gradual, end), palette, luma, contrast and a motion class. Turn them into shot length and pacing numbers.
- Reference ingest source block: rotation, vfr and aspect
Before you remake a reference, read the manifest source block: display_aspect, rotation, fps, vfr, codec and has_audio_track. What each field should change.
- Reference ingest uncertain[]: five kinds and what to do
The reference-ingest manifest lists five uncertain kinds, each with a suggested next step. Read the list, look again once per entry, and skip the rest.
Written by Sume