Extract on-screen text from a reference video with an API
Reference ingest reads on-screen text at source resolution, returns lines with boxes, spans and confidence, and flags low-confidence lines instead of guessing.

To pull the on-screen text out of a reference video, call POST /v1/reference-ingest. The manifest's text_tracks[] holds each line it read, with the text, a normalised box, a time span, a confidence and a role_hint. Nothing is corrected: a line below the confidence threshold is marked needs_verification and comes with a crop at the clip's native resolution.
Availability matters. Per the Reference ingest page read 2026-09-29, the endpoint is on in development and opt-in in production, so confirm it is listed for your workspace first.
Which languages does it read?
The OCR pass is set up for Korean and Latin text. ocr.languages defaults to ["ko", "en"]. It reads deduplicated frame states at source resolution and merges matches across frames into lines, so a caption that stays on screen for three seconds is one line, not ninety.
How do I tune it?
| Field | Meaning | Default |
|---|---|---|
ocr.languages | Languages to read | ["ko", "en"] |
ocr.fps | Text-state sampling rate, 0.5 to 2. OCR runs on deduplicated states, at most five frames | 1 |
ocr.min_confidence_attach_crop | Lines under this value are needs_verification and carry a crop | 0.85 |
curl -X POST https://api.sume.com/v1/reference-ingest \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ref-ocr-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/reference.mp4",
"ocr": { "languages": ["en"], "fps": 2, "min_confidence_attach_crop": 0.9 }
}'What should I do with a low-confidence line?
Read the crop, then, if you still need a pixel-level look, extract a frame at a manifest time with the video frames endpoint. The docs say to do this at most once per uncertain[] entry. Do not paste an unverified OCR line into a brief or an ad; a wrong price or claim copied from a reference is a real error, and the manifest tells you exactly which lines to check.
What can go wrong?
source_too_long_for_reference_ingest: the clip is over 300 seconds.source_no_video_stream: the file has no video.ffmpeg_fields_rejected: you sent a filter or codec key. Sume compiles every pass itself.400forsemantic: true, which is refused withreference_ingest_semantic_unavailableuntil that pass ships.
How do I get the crops back?
delivery.inline_strip and delivery.inline_crops control what the Sume Agent host attaches as inline images: the strip is on by default, and crops are none, low_confidence_only (default) or all. Over plain HTTP you read the manifest and its crops from the response. The hosted MCP reference_ingest tool returns the manifest as text.
Sources
Related posts
More in Developers
- Check whether a reference video is silent before adding music
Reference ingest reports audio.silent at a -60 LUFS gate, speech presence and beats, so you know whether to keep, replace or add a soundtrack before a remix.
- Remove background from image in Node.js (JavaScript API)
Remove an image background from Node.js: call a background-removal API with fetch on your server, poll the job, and save the transparent PNG.
- Remove background from image in Python with an API
Remove an image background in Python: POST the image URL with requests, poll the job, then save the transparent PNG. A full script and the price.
- Replicate API rate limits: 600 creates a minute, then 429
Replicate's API allows 600 prediction creates and 3,000 other requests per minute. Low credit and no card tighten it; over the limit you get a 429.
Written by Sume