Spot-check burned-in captions with video frames at STT word times

Check captions on a rendered video by pulling stills at word midpoints from a Sume STT result with POST /v1/video-frames, then compare text to speech.

4 min readSume
All posts

To spot-check burned-in captions without scrubbing the video by hand, take the word times from a Sume STT result, pick a few moments in the middle of words, and ask POST /v1/video-frames for stills at those times. Then look at each still and see whether the caption on screen matches the word being spoken. A dozen frames covers a one minute clip, and the whole check is one request and one read.

Video frames cuts stills from one Sume-hosted clip at the seconds you list in at[]. It does not edit or inspect the clip, and it is billed by its own Modal compute rather than a fixed price, so the docs do not give a per-frame number to quote.

Pick the times from word timings

An STT result has words[], each with word, start and end in seconds. For a spot check, take the midpoint of a handful of words: (start + end) / 2. Choose some near the start, some near the end, and a few after a pause, since drift tends to show up late. Skip entries whose type is spacing.

If you built the captions from your own script, the caption timing came from alignment against the STT words, so these same midpoints are the right places to look. If the on-screen words differ from the spoken word at that moment, you have either a timing offset or a script that differs from the speech.

Video frames request fields, from Sume docs read 2026-10-07
FieldValueNote
video_urlA media.sume.com clip of the workspaceThe API does not fetch the open internet; import first
at[]Seconds into the clipSend either at[] or fps, not both
formatjpeg (default) or pngpng avoids compression blur on small text
max_edge16 to 2160Long-edge clamp; omit to keep source size
Idempotency-KeyAny unique stringSo a retry does not queue a second extract

The request

Send the rendered, captioned clip as video_url, with the midpoints in at. Use png if the caption text is small, because jpeg artifacts can make thin glyphs hard to read. The submit always returns 202 with a request_id, and you read GET /v1/video-frames/{id} once the job is done.

curl -X POST https://api.sume.com/v1/video-frames \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: caption-check-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/captioned.mp4",
    "at": [1.2, 7.85, 14.3, 22.1, 31.6],
    "format": "png"
  }'
# poll GET /v1/jobs/$REQUEST_ID/status, then GET /v1/video-frames/$REQUEST_ID

What to look for in each still

Check three things in each still. The caption words should match the word being spoken at that second. The caption should not be cut off by the frame edge or covered by a platform's interface. The text should be legible against the background at that moment, since a bright scene can wash out a light caption.

The frames come back as durable images at the source size unless you clamp them, so you can attach them to a review ticket. If many frames show the same offset, fix the cause once: see captions out of sync with the audio for the usual culprits, such as chunk offsets when long audio was split.

When this is worth the cost

For one short video, watching it with the sound on is faster. A spot check pays off when you render many captioned clips from a template, because you can check the first of a batch, or a sample from each batch, and catch a systematic offset before the rest are published. Since the extract is billed by compute, keep the number of frames small and the max_edge modest when you only need to read text.

Restyling does not require re-running speech recognition: the captions endpoint accepts a source_caption_id, so a fix to style or placement can reuse the earlier timing. Re-check one frame afterward rather than the whole set. Keep the list of times you used with the job, so the same check can be repeated on the next render and compared against it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume