Read on-screen text in video: stills at 768 px, or raise max_edge
Video inspect returns 8 stills at 768 px by default. For small on-screen text, raise max_edge up to 2160 or ask for PNG. Settings, limits and cost.

If you need to read small on-screen text from a video, raise max_edge on Sume's video inspect. Stills default to a max_edge of 768 pixels in jpeg; the allowed range is 64 to 2160, and png is available through format. Ask for the larger size only on the frames you need, because each call returns at most 24 stills.
This ties to Google's agentic video understanding (Gemini API changelog, read 2026-10-02): a model that asks for frames on demand still needs frames sharp enough to read. On Sume the frames come from inspect, and reading them is up to your own model or a person.
Which settings matter?
The frames object takes exactly one of at or fps, never both.
| Field | Values | Default |
|---|---|---|
| frames | omitted, false, {at}, or {fps} | 8 mid-bin stills |
| at | 1 to 24 timestamps, each 0 or more | none |
| fps | above 0, up to 2, capped at 24 stills | none |
| max_edge | 64 to 2160 | 768 |
| format | jpeg or png | jpeg |
What does a text-reading request look like?
Pick the timestamps where text appears, set a large max_edge and use png for sharp edges.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: text-frames-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/ad.mp4", "frames": {"at": [2.5, 9, 14], "format": "png", "max_edge": 1600}}'What if the text is between keyframes?
Keep seek on precise, the default, which decodes to the exact instant. fast snaps to the keyframe at or before each instant and is for skimming.
Is it billed?
Probe and stills are unbilled; only the optional transcript is billed.
When is a bigger still not enough?
If the text is small, compressed or motion-blurred in the source, a larger still only enlarges the blur. Pick a frame where the text is still, and check the probe for the source resolution: a max_edge above the source size adds nothing. For a frame at the exact source size at one instant, Sume's video frames route is the better tool.
Also expect the limits of any reader. Whether text is legible depends on the reader as well as the image, so check a few frames by eye before you trust an automated pass.
How many frames should I ask for?
Ask for as few as will answer the question. Eight default stills at 768 pixels will show you the layout of a short ad; three precise timestamps at 1600 pixels will show you its price line. Stills come back as durable media.sume.com images, so you can reuse them without calling again.
How do I read the text afterwards?
Pass the stills to an agent model that accepts image input, or read them yourself. Sume returns the images and does not extract the text, so the accuracy of the reading depends on the reader you choose. Keep the still URL next to what was read so the claim can be checked later.
Sources
Related posts
More in Media tools
- Reference ingest OCR needs_verification: read the crop
Low-confidence on-screen text from reference ingest returns as needs_verification with a native crop. How the 0.85 default works and how to read the manifest.
- reference_ingest_semantic_unavailable: what to do instead
semantic: true is refused on reference ingest, and reference_ingest_unavailable means the media runtime lacks the function. The two errors and what to call.
- Restyle burned-in captions without paying for a second transcription
Pass source_caption_id to POST /v1/video-captions to re-burn the same video in another style, reusing its word timings. Billing is still one render.
- Revert an AI video edit: why Sume trims never touch your source
Descript moved its Revert button next to the AI response. In an API pipeline revert is free: each Sume edit returns a new artifact and the source stays put.
Written by Sume