Reference ingest OCR needs_verification: read the crop
Low-confidence on-screen text from reference ingest returns as needs_verification with a native crop. How the 0.85 default works and how to read the manifest.

When reference ingest is unsure about a line of on-screen text, it does not guess: the line is marked needs_verification and its native-resolution crop is returned so you, or the model, can read the pixels. The threshold is ocr.min_confidence_attach_crop, default 0.85.
This is the answer to "why is my OCR text wrong in the manifest": by design the text is never corrected, only flagged.
How does Sume read text in a reference video?
text_tracks[] comes from PP-OCRv5 (Korean and Latin) read at source resolution on deduplicated frame states, then merged across frames into lines. Each line has text, a normalised box, a span, a confidence, a persistence value, a card_id grouping and a role_hint.
Sampling is ocr.fps (0.5 to 2, default 1), and OCR runs on at most five deduplicated frames. So a caption that flashes for a fraction of a second between samples can be missed, which the docs do not hide: the manifest is evidence, and coverage says what ran.
What are the controls?
All are optional fields on POST /v1/reference-ingest:
| Field | Default | Effect |
|---|---|---|
ocr.languages | ["ko","en"] | Languages read |
ocr.fps | 1 | Text-state sampling rate, 0.5 to 2 |
ocr.min_confidence_attach_crop | 0.85 | Lines under it are needs_verification with a native crop |
delivery.inline_crops | low_confidence_only | What the agent host attaches as images: none, low_confidence_only, all |
How do I act on an uncertain line?
Treat uncertain[] as the only reason to look again. Read the attached crop, then request video_frames_create at a manifest time, at most once per entry. Do not loop on the same entry: if the crop is unreadable the answer is to say the text is unreadable.
On the Sume Agent host, sume-agent__reference_ingest returns the same manifest and attaches the strip plus up to four low-confidence crops as inline images, so the model sees facts and pixels in one turn.
curl -X POST https://api.sume.com/v1/reference-ingest \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ref-ocr-001" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_demo/ref.mp4",
"ocr":{"languages":["en"],"min_confidence_attach_crop":0.9}}'What should I change if too many lines are flagged?
A flood of needs_verification lines usually means stylised fonts, text over busy footage or small captions, not a bad call. Lowering ocr.min_confidence_attach_crop below 0.85 attaches fewer crops, which hides the problem rather than fixing it. Raising it toward 1 attaches more, which costs more image context in an agent turn.
A practical order of operations:
- Read
text_tracks[]for lines at or above the threshold and trust their spans andcard_idgroupings. - For each
uncertain[]entry, read the crop once. - If the crop settles it, write the corrected text in your own brief, not back into the manifest.
- If it does not, mark the text as unreadable and write new copy rather than guessing the original.
What does this not do?
It does not translate or correct text, and it does not read semantics: semantic: true is refused with reference_ingest_semantic_unavailable until the enrichment pass ships. Reference ingest is unbilled CPU work on the media runtime. Where a flag gates a route, the docs say so: these surfaces are listed where the matching environment flag allows them (development auto-on, production opt-in), so check tools_list or the OpenAPI before you build on one.
Where does the cost show up in usage?
The manifest itself is unbilled. Only speech.allow_billed_stt reserves the sume/video-inspect-1.0#transcript rate per minute of the hint (one minute when absent), transcribes only when the track is not silent and speech is found, and settles to what ran. Warnings such as stt_skipped_silent tell you when it was skipped.
If you run ingest as part of an agent workflow, the agent's own turns are billed separately from the ingest.
For budgeting, assume zero for the manifest. If you opt into the transcript, the Video inspect rate is $0.01 per audio minute, so a 300 second clip reserves at most about $0.05.
Sources
Related posts
More in Media tools
- reference_ingest_semantic_unavailable: what to do instead
semantic: true is refused on reference ingest, and reference_ingest_unavailable means the media runtime lacks the function. The two errors and what to call.
- Restyle burned-in captions without paying for a second transcription
Pass source_caption_id to POST /v1/video-captions to re-burn the same video in another style, reusing its word timings. Billing is still one render.
- Revert an AI video edit: why Sume trims never touch your source
Descript moved its Revert button next to the AI response. In an API pipeline revert is free: each Sume edit returns a new artifact and the source stays put.
- Runway Enhance Frame Rate: 10 target rates, 300 s, vs Sume
Runway Enhance Frame Rate takes clips up to 300 seconds at 1 credit per 2 seconds. Sume has no such model; its video-trim output.fps accepts 24, 25, 30 or 60.
Written by Sume