Caption a dubbed video: new job, not source_caption_id
source_caption_id reuses the first video's word timings, so a dub needs its own caption job on the dubbed file. How Sume handles it, and the cost.

Do not caption a dubbed video with source_caption_id. That field makes Sume reuse the source video and the word timings from an earlier caption job, so it is for restyling the same video, not for a version whose audio has changed. A dub is a new file with new audio, so submit its public HTTPS URL as video_url in a fresh job, with language set to the dub's language as a speech-to-text hint.
Everything here comes from the Video captions page of docs.sume.com, with timing and audio details from Audio detach.
What does source_caption_id actually reuse?
The docs say that passing source_caption_id instead of video_url makes Sume reuse that caption's source video and the word timings it already has, so no second speech-to-text runs. You may pass words alongside it only to correct the wording. Billing is unchanged: a restyle is still a render.
That is exactly right when the clip has not changed and you want a different look. It is wrong for a dub, because the words and times belong to the original language's speech.
Which caption route fits which situation?
Pick by what you already hold. The table lists the options the docs give for the caption body.
script_text, words, cues, and segments are mutually exclusive, so choose one per job.
| You have | Send | What happens |
|---|---|---|
| Only the dubbed MP4 | video_url, language | Speech-to-text on the dub, captions burned from its words |
| The dub plus its exact script | video_url, script_text | STT timings stay the source of truth; your wording is aligned to them |
| Per-phrase times you made | video_url, cues or segments | No speech-to-text; your text burns at your times |
| A restyle of the same clip | source_caption_id, style | Reuses source video and word timings; no second STT |
What can go wrong with script_text on a dub?
With script_text, alignment can fail with script_alignment_mismatch or script_alignment_failed, and the suggested next action is simplify_script_text_or_omit. Dubs raise the odds of a mismatch, because the voice may skip, merge, or reorder words relative to the script. Omit script_text to burn the speech-to-text wording instead, or fall back to cues when you know the times.
What does it cost, and what is the limit?
Each accepted standalone caption job reserves and captures $0.20 for a video up to 60 seconds under the current fixed estimate; confirm it in GET /v1/catalog. If your dubbed video is longer, the docs state no fixed rate for it here, so check the catalog before bulk runs.
The caption job needs a fetchable public HTTPS video_url. If you built the dub as a Sume-hosted file, use its media.sume.com URL. Pass a different Idempotency-Key per language and per attempt, and reuse a key only when you retry the same payload.
Sources
Related posts
More in Media tools
- caption_no_speech on a silent clip: burn Halloween text with cues
A silent AI clip fails video captions with caption_no_speech. Pass cues with text, start and end instead: a worked giveaway announcement on Sume at $0.20.
- Captions unreadable on busy footage: dim the clip, then burn
Busy B-roll can swallow burned captions. Sume can dim the clip with video-filter, then burn captions with colour overrides, for about $0.22 a clip.
- Change a video's background with a prompt: Gemini Omni on Sume
Google pitches Gemini Omni background changes by prompt. How to swap the background of your own clip with Sume's video edit mode: request and limits.
- Check character drift across AI shots with video_frames
Pull up to 24 evenly spaced stills from each AI clip with the unbilled video_frames route and compare faces and products before you stitch shots together.
Written by Sume