Caption a dubbed video: script_text and the Hangul style rule

After you dub a video, burn captions with script_text so the words match your script. Sume uses speech only for timing. A Korean script needs a Hangul style.

4 min readSume
All posts

To caption a dubbed video, call POST /v1/video-captions with the dubbed clip as video_url and your translated text as script_text. Sume keeps speech-to-text timing as the source of truth and aligns the burned words to your script, so the captions say what you wrote and not what a recognizer heard (video captions docs). If you omit script_text, the job transcribes the audio and burns that.

Rules that matter for a dub

  • The clip needs audible speech. A silent clip fails with caption_no_speech. For silent clips send authored cues with text, start and end instead.
  • language is only a speech-to-text hint. It does not pick the style or the font.
  • If you send Korean text to the Latin styles slam, punch or tiktok-green, the API returns a 400, caption_hangul_text_latin_style, rather than render empty boxes. Use a Hangul style such as black-outline, clip-wipe or korean-ad.
  • If you omit style, the text sets it: Latin text gets slam and Korean text gets black-outline.

Request

The dubbed clip must be on a public HTTPS URL. The fields below are all documented request fields.

{
  "video_url": "https://media.sume.com/artifacts/artf_demo/dub-ko.mp4",
  "script_text": "Your Korean script, one block of text.",
  "language": "ko",
  "style": "black-outline"
}

When timing and script disagree

The standalone captions endpoint fails hard on an alignment mismatch, according to the Sume API contract (API reference). A mismatch usually means the dub does not say what the script says, for instance when a line was re-recorded and the script was not updated. Fix the script or the audio, rather than loosening the match. If your dub came from Sume TTS with timestamps.words, you can also hand the timed words straight to the captions job, which skips speech-to-text.

Why check this now

Multilingual dubbing is easy to produce at scale. Microsoft's MAI-Voice-2.1 lists 23 languages from one voice (Microsoft AI, read 2026-10-04), so clips in languages nobody on the team reads are now normal. Captions sourced from the script give you a text you can have reviewed, which a recognizer's output is not.

A review routine

Keep the script file that produced the dub next to the captioned video. When a reviewer flags a line, edit the script, re-voice that line, and re-run the captions job with the new text. Because the burned words come from your script, a fix to the text is visible in the next render without anyone retyping subtitles by hand.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume