Fix one wrong word in burned-in captions without a re-transcribe
One misheard word in a burned-in caption? Re-burn with source_caption_id and a words list instead of running speech-to-text again. Request body included.

The short answer
Send a new POST /v1/video-captions with source_caption_id set to the earlier caption job and a words list that carries the corrected wording. The Video captions docs say Sume then reuses that caption's source video and the word timings it already has, so no second speech-to-text runs, and that you pass words alongside the id only to correct the wording. Billing is unchanged: a restyle is still a render, at the $0.20 per job estimate for videos up to 60 seconds.
Why not just resend the video
Sending the video again with a fresh video_url runs speech-to-text again, and it can mishear the same word the same way. Passing script_text is the other documented fix: Sume keeps speech-to-text timings and aligns your wording to them, but alignment can fail with typed errors, script_alignment_mismatch or script_alignment_failed, whose suggested next action is simplify_script_text_or_omit. The script alignment fix post covers that route.
The source_caption_id route sidesteps both. It does not realign anything; it burns the words you give at the times you give.
The request
Each entry in words has text, start and end, in seconds. The docs state that script_text, words, cues and segments are mutually exclusive, so do not combine words with script_text. They do not say whether a partial list is accepted, and raw transcripts are not part of the public contract, so keep the word timings from your own earlier request or build the full list for the clip yourself. Treat the sample below as a pattern, not a recovered transcript.
import json
words = [
{"text": "Meet", "start": 0.0, "end": 0.3},
{"text": "Sume", "start": 0.3, "end": 0.7},
{"text": "Avatar", "start": 0.7, "end": 1.2},
]
fix = {"Sume": "Sume"}
words = [{**w, "text": fix.get(w["text"], w["text"])} for w in words]
body = {
"source_caption_id": "vc_123",
"style": "black-outline",
"words": words,
}
print(json.dumps(body, indent=2))Which fix to use
The right route depends on what is wrong.
| Situation | Route | Runs speech-to-text again |
|---|---|---|
| First render, spelling known in advance | script_text | Yes, timings from speech-to-text |
| One word wrong after a render | source_caption_id plus words | No |
| New look, wording fine | source_caption_id and a new style | No |
| Silent clip or fully authored copy | cues, segments or words with video_url | No |
After you submit
The result follows the normal job lifecycle in the Jobs and results docs: poll the status, read the result once it completes, and reuse the same Idempotency-Key on a retry so a dropped connection does not render twice. Check the new video for the corrected word and for the surrounding words, because the style you pick controls how many words sit on a card.
If the docs left a detail open for your case, such as a partial words list, test it on one short clip before correcting a batch. The restyle post has the basic id flow if you have not used it yet.
Sources
Related posts
More in Media tools
- Cover or contain: fit a 16:9 clip into a vertical stack
Sume timeline compose takes video_fit cover, contain or stretch. Work out what each does to a 16:9 clip in the 1080x1920 default stack, then check one frame.
- Cut a clip into 30-second chunks for 30-second post-processing caps
Runway's SDR-to-HDR model and the Magnific upscaler take 30 seconds at most. Cut any clip into 30-second pieces with Sume video trim at $0.02 per cut.
- Douyin ad first frame: black at most 60%, check with stills
Douyin in-feed ads need a first frame that is at most 60% black. Pull stills from your render with Sume video inspect and measure the first frame before upload.
- Edits carousels: four matching images per call
Instagram Edits now supports carousels, per secondary reports. Sume image models with an n range of 1-4 return up to four matching images per request.
Written by Sume