reference_ingest OCR: needs_verification crops under 0.85 confidence
reference_ingest never corrects OCR text. Lines under 0.85 confidence come back as needs_verification with a native-resolution crop to check.

When reference_ingest reads text off a clip and is not sure, it says so. Any OCR line under the confidence threshold, 0.85 by default, comes back marked needs_verification with a crop at native resolution, so you check the pixels rather than trust a guess. The tool corrects nothing. That matters when the text is a price, a claim or a brand name.
What the text read looks like
The manifest's text_tracks[] come from PP-OCRv5 (Korean and Latin) at source resolution, run on deduplicated frame states and merged across frames into lines. Each line carries text, a normalized box, a time span, a confidence, a persistence value, a card_id group and a role_hint. OCR runs on at most five sampled frames, so text_tracks[] is not caption coverage: a caption that appears only between samples can be missing.
The knobs
Three request fields shape the read. They are optional, and they all live under ocr or delivery. Keep the defaults until a clip proves they are wrong.
| Field | Default | Effect |
|---|---|---|
| ocr.languages | [ko, en] | Languages read |
| ocr.fps | 1 (0.5 to 2) | Sample rate for text states; at most five frames |
| ocr.min_confidence_attach_crop | 0.85 (0 to 1) | Lines below it are needs_verification with a native crop |
| delivery.inline_crops | low_confidence_only | none, low_confidence_only or all (agent host) |
| delivery.inline_strip | On | Labeled overview strip as an image |
Check, do not edit
The manifest's uncertain[] list is the only reason to look at a frame again. For each entry, read the attached crop, and if you still need a clean frame, call video_frames at the manifest time, at most once per entry. On the Sume Agent host the strip plus up to four low-confidence crops arrive as inline images, so the model reads facts and pixels in one turn. Over hosted MCP the text result lists them, and you fetch the crop yourself.
Raise ocr.min_confidence_attach_crop toward 1 when exact wording matters, such as a legal line or a price. You get more crops to check and a higher chance of catching a misread. Lower it when you only want the gist.
What goes wrong with price and brand text
Small stylized type, text over motion, and low-contrast overlays are where OCR gives a confident-looking wrong string. Never copy an OCR price straight into a script. Treat the line as a pointer, open the crop, and type the value yourself. If the same card_id appears across several spans, compare them; a number that changes between reads of the same card deserves a human look. Log each crop's job id with the decision you made about it, so a later reviewer can see what was checked and by whom. A note that says 'verified by eye at 0:14' is worth more than a clean-looking string.
Sources
Related posts
More in Developers
- Reference-to-video with 3 images, 6 seconds: cost by Sume model
A 6-second reference-to-video clip with three images costs $0.45 on H3 at 768p up to $3.47 on Seedance 2.5 at 720p on Sume. Eight rows priced.
- Resume a video chain after a worker crash from stored job ids
Your worker died between generate and trim. Read the stored job id from /v1/jobs, do not resubmit paid work, and continue from the first step with no result.
- Retry a Kling motion control submit without paying twice
All four Sume motion-control routes take an Idempotency-Key. Resend the same body and key after a timeout and you get the original job, not a second reserve.
- Retry a Sume video submit safely: Idempotency-Key and the 409 conflict
Send the same Idempotency-Key and the same body to retry a Sume video submit without a second charge. A different body under the same key returns 409.
Written by Sume