Reference ingest coverage: how many frames OCR read

The reference-ingest coverage block reports frames decoded, frames OCR read and the OCR rate. Use it to decide when on-screen text needs a second look.

5 min readSume
All posts

The coverage block in a Sume reference-ingest manifest tells you how much of the clip the text reader actually saw: frames_decoded, frames_ocr, ocr_fps, ocr_on_cut_frames and whole_duration. If frames_ocr is small next to the clip length, a title card that flashed between samples may not be in text_tracks[], and you should look at the strip or pull frames before you say the reference has no text.

Episode formats lean on short title cards. YouTube's September 23 post describes Shorts series as seasons and episodes with custom thumbnails (YouTube Blog, read 2026-10-04), so a competitor episode you ingest may carry a one-second card.

What the sampling does

Per the reference ingest docs, ocr.fps is 0.5 to 2 with a default of 1, and OCR runs on deduplicated text states, at most five frames. The docs are explicit that text_tracks[] are OCR hits on at most five sampled frames, not caption coverage. So for a 30-second clip at the default rate there are about 30 candidate instants, and at most five frames are read.

coverage fields as listed in the contract and OpenAPI schema (read 2026-10-04)
FieldTypeUse
frames_decodednumberHow many frames the pass decoded
frames_ocrnumberHow many frames OCR actually read
ocr_fpsnumberThe sampling rate used, 0.5 to 2
ocr_on_cut_framesbooleanWhether OCR also ran on cut frames
whole_durationbooleanWhether the read spanned the whole clip

A three-line verdict

Treat an empty text_tracks[] as unproven when whole_duration is false or frames_ocr is 5 on a text-heavy clip. Raising ocr.fps toward 2 samples more instants but does not lift the five-frame ceiling the docs describe, so read the strip or extract frames at the times you care about.

coverage = {
    "frames_decoded": 61,
    "frames_ocr": 5,
    "ocr_fps": 1,
    "ocr_on_cut_frames": True,
    "whole_duration": True,
}
text_tracks = []

if not coverage["whole_duration"]:
    verdict = "partial read: do not trust an empty text list"
elif coverage["frames_ocr"] >= 5 and not text_tracks:
    verdict = "five frames read, none had text: check the strip"
elif not text_tracks:
    verdict = "whole clip read, no text found"
else:
    verdict = f"{len(text_tracks)} text tracks"
print(verdict)

Reference ingest is dest first and production opt-in. When a line is read at low confidence it comes back as needs_verification with a native crop; see OCR crops. Extra frames come from video frames.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume