English to Korean captions: style follows the text, not language
Localising a video to Korean? Sume's caption style follows the wording, and a Latin style on Hangul text is a 400. Pick styles per locale before a batch.

In a localisation batch, do not hard-code one caption style for every language. The language field on Video captions only hints speech-to-text and never selects the style or font; an omitted style follows the wording, and naming a Latin style with Hangul text returns 400 caption_hangul_text_latin_style.
What decides the style when I leave it out?
Omit style and the wording decides: Latin copy gets slam, Korean copy gets black-outline. A style you name is rendered as named. The doc is explicit that language never picks the style or the font, and that it is only a speech-to-text hint (ko, en, and so on; omit it for automatic detection).
What fails when a batch reuses one style?
slam, punch and tiktok-green draw in Latin display faces with no Hangul glyphs. Sending Korean text to one returns 400 instead of burning tofu boxes, and a Hangul font with a Latin style fails as caption_font_requires_hangul_style. A batch that hard-codes slam will fail every Korean row.
korean-ad is not what an omitted style resolves to. Ask for it explicitly when you want the ad karaoke look, and pair it with language: "ko" for Korean speech.
| Target text | Style omitted | Named styles that work | Font |
|---|---|---|---|
| English or other Latin | slam | slam, punch, tiktok-green | Keep the style's own face |
| Korean | black-outline | black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis, korean-ad | Optional Hangul face |
How do you localise a finished video?
Sume's caption call burns the text you give it; this post assumes you have the translated lines already. Pass them as cues with text, start and end in seconds, reusing the timings of the original sentences from Video inspect sentence segments. Cues skip speech-to-text, so the language hint does not matter, but the wording still decides the default style.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: intro-ko-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/intro.mp4",
"style": "black-outline",
"cues": [
{"text": "오늘은 첫 번째 강의입니다.", "start": 0.0, "end": 2.4}
]
}'What else should I check before a bulk run?
Branch the style per target language in your script, or leave it out and accept the defaults. Korean lines are often shorter or longer than the English ones, so check the line fits in the cue time. Each standalone job is $0.20 for videos up to 60 seconds under the current estimate. If you restyle later, source_caption_id re-burns a finished caption job without a second speech-to-text run.
How do you handle mixed Korean and English lines?
A cue list that mixes scripts is the awkward case. The docs say an omitted style follows the wording and that Korean copy resolves to black-outline. If a line has English brand names inside Korean text, use a Hangul style such as black-outline, and look at the render: Hangul faces are the ones the doc describes, and Latin glyph coverage inside them is not documented, so verify brand names visually.
Fonts follow the same rule. Setting a Hangul font with a Latin style fails, and any name outside the documented list is rejected rather than substituted, so a caption never falls back to a face you did not ask for. All listed faces ship with the renderer under the SIL Open Font License 1.1.
What is a quick decision rule?
If your batch covers one language per run, name the style you want and keep it for the run. If one run mixes Latin and Hangul targets, build the request per target: Latin rows get slam or nothing, Korean rows get black-outline or one of the Hangul identities. Never share a single style value across both unless it is omitted.
Test one row of each language first. A single request costs $0.20 for a clip up to 60 seconds, and a wrong style on a Korean row is a 400 you see immediately, before the bulk run starts.
What does this not cover?
Sume does not translate in this call. The caption job burns the text you give it, so the translation is your input.
Sources
Related posts
More in Media tools
- Extract audio from an AI avatar video as a WAV with one API call
Sume audio detach pulls the track of one hosted video into a WAV or MP3 for $0.01. The request, the 16 kHz mono option, the caps and the no-audio error.
- Extract audio from a video for transcription: 16 kHz mono wav
Detach the audio from a Sume-hosted video as 16 kHz mono wav, then send it to speech-to-text. Costs, caps and the errors to expect.
- Facebook Reels audio spec: AAC-LC, stereo, 48 kHz, 128 kbps+
Facebook Reels wants AAC Low Complexity, stereo, 48 kHz and 128 kbps or more. What a Sume probe shows, what trim keeps, and what you must verify yourself.
- FFmpeg drawtext on a hosted API: why Sume filters refuse it, and cues
Sume video-filter refuses drawtext and subtitles because they read files. To burn text onto a clip, send cues to video-captions: text, start and end seconds.
Written by Sume