Captions language field: a speech hint, not a style or font

On Sume video captions, language only hints speech-to-text. It never picks the style or font, so Korean needs a Hangul style such as black-outline.

4 min readSume
All posts

The language field on a Sume caption job only tells speech-to-text which language to expect. It never selects the look or the font. If you set language: "ko" and also pick a Latin style such as slam, the API answers 400 with caption_hangul_text_latin_style for Korean text, because those styles have no Hangul glyphs.

Rules from the Video captions docs (read 2026-10-06).

What does each field do?

style picks the look and the motion, design changes that look for one request, font picks the face for a Hangul style, and language is only a speech hint. If you omit language, speech-to-text finds the language automatically.

Caption fields and what they control (read 2026-10-06)
FieldControlsDefault when omitted
styleLook and motionslam for Latin text, black-outline for Korean text
fontHangul face for a Hangul styleThe style's own face
designPer-request overrides of the styleThe style's tokens
languageSpeech-to-text hintAutomatic detection

Which styles take Korean?

Use black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis or the ad look korean-ad. slam, punch and tiktok-green are Latin display faces and are rejected for Hangul text. A font value that is a Hangul face with a Latin style returns caption_font_requires_hangul_style.

Why is rejection the better outcome?

The docs note that a clip full of tofu boxes would cost the same as a good one, so Sume returns an error before the render and does not switch your style.

Sources

More in Media tools

All Media tools posts

Written by Sume