Localize One Ad to Korean: Burned-In Captions With Hangul Styles

Add Korean captions to an ad: why Latin styles return a 400, which Hangul style to pick, how a restyle skips a second transcription, and what a job costs.

5 min readSume
All posts

The short answer

Send the video to POST /v1/video-captions with language: "ko" and a Hangul style such as korean-ad or black-outline. If you send Korean text to the Latin styles slam, punch or tiktok-green, the API returns 400 caption_hangul_text_latin_style, because those fonts have no Hangul glyphs and would draw empty boxes. A standalone caption job is $0.20 for a video up to 60 seconds, per the docs page read on 2026-10-09.

Pick the style

Docs list six Hangul identities for Korean speech and an ad look. If you do not send style, Korean text resolves to black-outline. korean-ad has to be chosen on purpose; it is CapCut-style karaoke, one short phrase at a time in the lower third, with the spoken word at a heavy weight. The language field is only a speech-to-text hint and does not choose the style or the font.

Hangul caption styles (Sume docs, read 2026-10-09)
StyleLook
korean-adAd karaoke: weight shift with an accent on the spoken word
black-outlineWhite fill on thick black outline, center of frame; the default for Korean
weight-shiftPhrase cards; the spoken word gets the heavy weight
highlightAn accent block moves behind the spoken word
pill-karaokeDark pill card; color changes with the voice
clip-wipeWords wipe in left to right; clearest at small phone sizes
editorial-emphasisTwo-line left-aligned card with a large last word

Fonts, scripts and restyles

The font field only applies to Hangul styles, and the faces are listed in the docs: Pretendard for most identities, Do Hyeon for black-outline and clip-wipe, and a longer list including Black Han Sans, Jua and Gmarket Sans. Sending a Hangul face to a Latin style returns 400 caption_font_requires_hangul_style. The weight-shift and korean-ad motion animates the weight axis, which only Pretendard has.

If you have the translated script, send it as script_text. Sume keeps the speech-to-text word timings as the clock and aligns your text to them; a mismatch gives script_alignment_mismatch, and the suggested next step is to simplify the script or omit it. To restyle a finished render, send source_caption_id and a new style: Sume reuses the stored word timings, so speech-to-text does not run again, and the price is still one render.

A workable routine for a campaign is to caption once in black-outline, show the client, then restyle with source_caption_id into korean-ad or clip-wipe if they want more motion. The restyle is still a render and costs the same $0.20 for a video up to 60 seconds, but it skips the second speech-to-text run. Check the live price in GET /v1/catalog before a long campaign.

  • Docs describe the caption job as transcribing the speech in the video; they do not describe translation. If the audio is English, supply Korean text as cues or words, or make a Korean voice track first.
  • Use cues (text, start, end) for a silent clip; speech-to-captions on a silent clip fails with caption_no_speech.
  • Have a Korean speaker review the final text; auto-timing is not a proofread.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume