Caption a dubbed video: script_text and the Hangul style rule
After you dub a video, burn captions with script_text so the words match your script. Sume uses speech only for timing. A Korean script needs a Hangul style.

To caption a dubbed video, call POST /v1/video-captions with the dubbed clip as video_url and your translated text as script_text. Sume keeps speech-to-text timing as the source of truth and aligns the burned words to your script, so the captions say what you wrote and not what a recognizer heard (video captions docs). If you omit script_text, the job transcribes the audio and burns that.
Rules that matter for a dub
- The clip needs audible speech. A silent clip fails with
caption_no_speech. For silent clips send authoredcueswithtext,startandendinstead. languageis only a speech-to-text hint. It does not pick the style or the font.- If you send Korean text to the Latin styles
slam,punchortiktok-green, the API returns a 400,caption_hangul_text_latin_style, rather than render empty boxes. Use a Hangul style such asblack-outline,clip-wipeorkorean-ad. - If you omit
style, the text sets it: Latin text getsslamand Korean text getsblack-outline.
Request
The dubbed clip must be on a public HTTPS URL. The fields below are all documented request fields.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/dub-ko.mp4",
"script_text": "Your Korean script, one block of text.",
"language": "ko",
"style": "black-outline"
}When timing and script disagree
The standalone captions endpoint fails hard on an alignment mismatch, according to the Sume API contract (API reference). A mismatch usually means the dub does not say what the script says, for instance when a line was re-recorded and the script was not updated. Fix the script or the audio, rather than loosening the match. If your dub came from Sume TTS with timestamps.words, you can also hand the timed words straight to the captions job, which skips speech-to-text.
Why check this now
Multilingual dubbing is easy to produce at scale. Microsoft's MAI-Voice-2.1 lists 23 languages from one voice (Microsoft AI, read 2026-10-04), so clips in languages nobody on the team reads are now normal. Captions sourced from the script give you a text you can have reviewed, which a recognizer's output is not.
A review routine
Keep the script file that produced the dub next to the captioned video. When a reviewer flags a line, edit the script, re-voice that line, and re-run the captions job with the new text. Because the burned words come from your script, a fix to the text is visible in the next render without anyone retyping subtitles by hand.
Sources
Related posts
More in Media tools
- Caption a silent AI video: fixing caption_no_speech
A silent clip fails POST /v1/video-captions with caption_no_speech. Send cues with text, start and end to burn authored captions without speech-to-text.
- Caption a silent Seedance 2.5 clip with authored cues
A clip with no speech fails Sume caption jobs as caption_no_speech. Pass cues with text, start and end seconds instead; $0.20 per job up to 60 seconds.
- Change the caption highlight colour without changing the style
Cheap transcription made captions routine; brand colour is what is left. One design field changes the spoken-word colour and keeps everything else in the style.
- Turn a music composition plan into a Sume time-range prompt (Python)
ElevenLabs music_v2_5 plans allow 6,132 characters in up to 30 lines. Sume's Music prompt takes 5000 characters. A Python converter for time ranges.
Written by Sume