Read on-screen text in video: stills at 768 px, or raise max_edge

Video inspect returns 8 stills at 768 px by default. For small on-screen text, raise max_edge up to 2160 or ask for PNG. Settings, limits and cost.

4 min readSume
All posts

If you need to read small on-screen text from a video, raise max_edge on Sume's video inspect. Stills default to a max_edge of 768 pixels in jpeg; the allowed range is 64 to 2160, and png is available through format. Ask for the larger size only on the frames you need, because each call returns at most 24 stills.

This ties to Google's agentic video understanding (Gemini API changelog, read 2026-10-02): a model that asks for frames on demand still needs frames sharp enough to read. On Sume the frames come from inspect, and reading them is up to your own model or a person.

Which settings matter?

The frames object takes exactly one of at or fps, never both.

Frame settings. Sume Video inspect docs, read 2026-10-02.
FieldValuesDefault
framesomitted, false, {at}, or {fps}8 mid-bin stills
at1 to 24 timestamps, each 0 or morenone
fpsabove 0, up to 2, capped at 24 stillsnone
max_edge64 to 2160768
formatjpeg or pngjpeg

What does a text-reading request look like?

Pick the timestamps where text appears, set a large max_edge and use png for sharp edges.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: text-frames-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/ad.mp4", "frames": {"at": [2.5, 9, 14], "format": "png", "max_edge": 1600}}'

What if the text is between keyframes?

Keep seek on precise, the default, which decodes to the exact instant. fast snaps to the keyframe at or before each instant and is for skimming.

Is it billed?

Probe and stills are unbilled; only the optional transcript is billed.

When is a bigger still not enough?

If the text is small, compressed or motion-blurred in the source, a larger still only enlarges the blur. Pick a frame where the text is still, and check the probe for the source resolution: a max_edge above the source size adds nothing. For a frame at the exact source size at one instant, Sume's video frames route is the better tool.

Also expect the limits of any reader. Whether text is legible depends on the reader as well as the image, so check a few frames by eye before you trust an automated pass.

How many frames should I ask for?

Ask for as few as will answer the question. Eight default stills at 768 pixels will show you the layout of a short ad; three precise timestamps at 1600 pixels will show you its price line. Stills come back as durable media.sume.com images, so you can reuse them without calling again.

How do I read the text afterwards?

Pass the stills to an agent model that accepts image input, or read them yourself. Sume returns the images and does not extract the text, so the accuracy of the reading depends on the reader you choose. Keep the still URL next to what was read so the claim can be checked later.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume