Fix one wrong word in burned-in captions without a re-transcribe

One misheard word in a burned-in caption? Re-burn with source_caption_id and a words list instead of running speech-to-text again. Request body included.

6 min readSume
All posts

The short answer

Send a new POST /v1/video-captions with source_caption_id set to the earlier caption job and a words list that carries the corrected wording. The Video captions docs say Sume then reuses that caption's source video and the word timings it already has, so no second speech-to-text runs, and that you pass words alongside the id only to correct the wording. Billing is unchanged: a restyle is still a render, at the $0.20 per job estimate for videos up to 60 seconds.

Why not just resend the video

Sending the video again with a fresh video_url runs speech-to-text again, and it can mishear the same word the same way. Passing script_text is the other documented fix: Sume keeps speech-to-text timings and aligns your wording to them, but alignment can fail with typed errors, script_alignment_mismatch or script_alignment_failed, whose suggested next action is simplify_script_text_or_omit. The script alignment fix post covers that route.

The source_caption_id route sidesteps both. It does not realign anything; it burns the words you give at the times you give.

The request

Each entry in words has text, start and end, in seconds. The docs state that script_text, words, cues and segments are mutually exclusive, so do not combine words with script_text. They do not say whether a partial list is accepted, and raw transcripts are not part of the public contract, so keep the word timings from your own earlier request or build the full list for the clip yourself. Treat the sample below as a pattern, not a recovered transcript.

import json

words = [
    {"text": "Meet", "start": 0.0, "end": 0.3},
    {"text": "Sume", "start": 0.3, "end": 0.7},
    {"text": "Avatar", "start": 0.7, "end": 1.2},
]
fix = {"Sume": "Sume"}
words = [{**w, "text": fix.get(w["text"], w["text"])} for w in words]

body = {
    "source_caption_id": "vc_123",
    "style": "black-outline",
    "words": words,
}
print(json.dumps(body, indent=2))

Which fix to use

The right route depends on what is wrong.

Ways to correct burned-in caption wording, from the Sume video captions docs (read 2026-10-04)
SituationRouteRuns speech-to-text again
First render, spelling known in advancescript_textYes, timings from speech-to-text
One word wrong after a rendersource_caption_id plus wordsNo
New look, wording finesource_caption_id and a new styleNo
Silent clip or fully authored copycues, segments or words with video_urlNo

After you submit

The result follows the normal job lifecycle in the Jobs and results docs: poll the status, read the result once it completes, and reuse the same Idempotency-Key on a retry so a dropped connection does not render twice. Check the new video for the corrected word and for the surrounding words, because the style you pick controls how many words sit on a card.

If the docs left a detail open for your case, such as a partial words list, test it on one short clip before correcting a batch. The restyle post has the basic id flow if you have not used it yet.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume