Spot-check burned-in captions with video frames at STT word times
Check captions on a rendered video by pulling stills at word midpoints from a Sume STT result with POST /v1/video-frames, then compare text to speech.

To spot-check burned-in captions without scrubbing the video by hand, take the word times from a Sume STT result, pick a few moments in the middle of words, and ask POST /v1/video-frames for stills at those times. Then look at each still and see whether the caption on screen matches the word being spoken. A dozen frames covers a one minute clip, and the whole check is one request and one read.
Video frames cuts stills from one Sume-hosted clip at the seconds you list in at[]. It does not edit or inspect the clip, and it is billed by its own Modal compute rather than a fixed price, so the docs do not give a per-frame number to quote.
Pick the times from word timings
An STT result has words[], each with word, start and end in seconds. For a spot check, take the midpoint of a handful of words: (start + end) / 2. Choose some near the start, some near the end, and a few after a pause, since drift tends to show up late. Skip entries whose type is spacing.
If you built the captions from your own script, the caption timing came from alignment against the STT words, so these same midpoints are the right places to look. If the on-screen words differ from the spoken word at that moment, you have either a timing offset or a script that differs from the speech.
| Field | Value | Note |
|---|---|---|
| video_url | A media.sume.com clip of the workspace | The API does not fetch the open internet; import first |
| at[] | Seconds into the clip | Send either at[] or fps, not both |
| format | jpeg (default) or png | png avoids compression blur on small text |
| max_edge | 16 to 2160 | Long-edge clamp; omit to keep source size |
| Idempotency-Key | Any unique string | So a retry does not queue a second extract |
The request
Send the rendered, captioned clip as video_url, with the midpoints in at. Use png if the caption text is small, because jpeg artifacts can make thin glyphs hard to read. The submit always returns 202 with a request_id, and you read GET /v1/video-frames/{id} once the job is done.
curl -X POST https://api.sume.com/v1/video-frames \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-check-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/captioned.mp4",
"at": [1.2, 7.85, 14.3, 22.1, 31.6],
"format": "png"
}'
# poll GET /v1/jobs/$REQUEST_ID/status, then GET /v1/video-frames/$REQUEST_IDWhat to look for in each still
Check three things in each still. The caption words should match the word being spoken at that second. The caption should not be cut off by the frame edge or covered by a platform's interface. The text should be legible against the background at that moment, since a bright scene can wash out a light caption.
The frames come back as durable images at the source size unless you clamp them, so you can attach them to a review ticket. If many frames show the same offset, fix the cause once: see captions out of sync with the audio for the usual culprits, such as chunk offsets when long audio was split.
When this is worth the cost
For one short video, watching it with the sound on is faster. A spot check pays off when you render many captioned clips from a template, because you can check the first of a batch, or a sample from each batch, and catch a systematic offset before the rest are published. Since the extract is billed by compute, keep the number of frames small and the max_edge modest when you only need to read text.
Restyling does not require re-running speech recognition: the captions endpoint accepts a source_caption_id, so a fix to style or placement can reuse the earlier timing. Re-check one frame afterward rather than the whole set. Keep the list of times you used with the job, so the same check can be repeated on the next render and compared against it.
Sources
Related posts
More in Developers
- Start a Sume render from a serverless function: submit, save, 202
A function must not wait for a video. Submit with mode webhook and a stable Idempotency-Key, save the status URL, return 202, and let the signed webhook finish.
- A Sume job looks stuck: wait, cancel or poll the events?
Read status, then events. Queued and processing mean wait, cancel only works before generation starts, and a client timeout never cancels the job.
- Sume timeouts in one table: 30 s, 55 s, 10 s, 90 minutes
Every wait in the Sume API has its own number: sync 30 s, jobs_wait 55 s, webhook attempts 10 s, SDK helpers 10 and 20 minutes, Format runs 90 minutes.
- Can a Sume webhook arrive twice? Build an idempotent receiver
Sume retries failed webhook deliveries up to 10 times and Redeliver replays a real event, so one terminal event can reach you twice. Dedupe on job_id or run_id.
Written by Sume