Dubbed video: burn target-language captions with script_text
After dubbing into Spanish or German, caption the new audio with your translated script as script_text so burned words match your approved translation.

After you dub a video into Spanish or German, burn captions from the translation you already approved, not from a fresh transcription. Send the dubbed MP4 to Sume's video captions endpoint with your translated script as script_text. Speech-to-text still supplies the timings, and Sume aligns your wording onto them, so brand names and product terms appear exactly as written. The job is $0.20 for clips up to 60 seconds.
Why not let the transcriber write it
A transcriber hears the dubbed audio and writes what it heard, which may differ from your approved translation: a different spelling of a product name, a number written as digits or words, a dropped filler. If your translation went through review, that review is the source of truth. script_text keeps it that way.
| Source | Wording comes from | Timing comes from | Risk |
|---|---|---|---|
No script_text | Speech-to-text of the dub | Speech-to-text | Misheard names |
script_text | Your approved script | Speech-to-text word timings | Alignment can fail |
words | Your words with times | Your times | You must have the timings |
cues | Your phrases with times | Your times | No speech needed; manual timing |
The request
Set language to the dub language so recognition expects it. Name a style that fits the script: the Latin styles for Spanish or German, a Hangul-safe style for Korean.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: dub-es-caption-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/dub-es.mp4",
"language": "es",
"style": "slam",
"script_text": "Conoce la nueva coleccion de verano."
}'When alignment fails
If the script and the speech differ too much, the job fails with script_alignment_mismatch or script_alignment_failed, and the suggested next step is to simplify the script or omit it. The usual cause in dubbing is that the dub speaks a slightly different sentence than the script, because the voice was given an edited version. Make the script you caption the same text you sent to text-to-speech, character for character, which should reduce alignment failures. The captions docs list the errors.
A note on voices and languages
Microsoft's MAI-Voice-2.1 page lists 23 languages (read 2026-10-04), and MAI-Transcribe-2-Streaming is described as covering 60 languages. Whatever produced your dub, the caption step here only needs a public HTTPS video URL, the language and the script. See the dubbing pipeline walkthrough for the steps before this one, and jobs and results for polling.
Sources
Related posts
More in Use cases
- CAWG identity assertion 1.2: who made it, not whether it is AI
A CAWG identity assertion says which named actor stands behind an asset. It does not say the asset is AI-generated. What Sume stores that you can pair with it.
- CAWG training and data mining 1.1: notAllowed vs constrained
The CAWG training-and-data-mining assertion has four entries, each allowed, notAllowed or constrained. Constrained with no extra info counts as notAllowed.
- Children's story clip from one illustration: Seedance 2.5 or 2.0
Animate one storybook illustration with a Seedance reference image on Sume. A 15 s page is $5.67 on Seedance 2.0 and $8.67 on Seedance 2.5.
- Christmas ad music: an instrumental Lyria brief that sells
A Christmas ad needs music that is festive without vocals under the voiceover. Brief it in one prompt, set the bed under speech and fade it out, with Sume.
Written by Sume