Caption a dubbed video: new job, not source_caption_id

source_caption_id reuses the first video's word timings, so a dub needs its own caption job on the dubbed file. How Sume handles it, and the cost.

4 min readSume
All posts

Do not caption a dubbed video with source_caption_id. That field makes Sume reuse the source video and the word timings from an earlier caption job, so it is for restyling the same video, not for a version whose audio has changed. A dub is a new file with new audio, so submit its public HTTPS URL as video_url in a fresh job, with language set to the dub's language as a speech-to-text hint.

Everything here comes from the Video captions page of docs.sume.com, with timing and audio details from Audio detach.

What does source_caption_id actually reuse?

The docs say that passing source_caption_id instead of video_url makes Sume reuse that caption's source video and the word timings it already has, so no second speech-to-text runs. You may pass words alongside it only to correct the wording. Billing is unchanged: a restyle is still a render.

That is exactly right when the clip has not changed and you want a different look. It is wrong for a dub, because the words and times belong to the original language's speech.

Which caption route fits which situation?

Pick by what you already hold. The table lists the options the docs give for the caption body.

script_text, words, cues, and segments are mutually exclusive, so choose one per job.

Caption inputs for a dubbed file, from docs.sume.com Video captions, read 2026-10-02.
You haveSendWhat happens
Only the dubbed MP4video_url, languageSpeech-to-text on the dub, captions burned from its words
The dub plus its exact scriptvideo_url, script_textSTT timings stay the source of truth; your wording is aligned to them
Per-phrase times you madevideo_url, cues or segmentsNo speech-to-text; your text burns at your times
A restyle of the same clipsource_caption_id, styleReuses source video and word timings; no second STT

What can go wrong with script_text on a dub?

With script_text, alignment can fail with script_alignment_mismatch or script_alignment_failed, and the suggested next action is simplify_script_text_or_omit. Dubs raise the odds of a mismatch, because the voice may skip, merge, or reorder words relative to the script. Omit script_text to burn the speech-to-text wording instead, or fall back to cues when you know the times.

What does it cost, and what is the limit?

Each accepted standalone caption job reserves and captures $0.20 for a video up to 60 seconds under the current fixed estimate; confirm it in GET /v1/catalog. If your dubbed video is longer, the docs state no fixed rate for it here, so check the catalog before bulk runs.

The caption job needs a fetchable public HTTPS video_url. If you built the dub as a Sume-hosted file, use its media.sume.com URL. Pass a different Idempotency-Key per language and per attempt, and reuse a key only when you retry the same payload.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume