Captions language field: a speech hint, not a style or font
On Sume video captions, language only hints speech-to-text. It never picks the style or font, so Korean needs a Hangul style such as black-outline.

The language field on a Sume caption job only tells speech-to-text which language to expect. It never selects the look or the font. If you set language: "ko" and also pick a Latin style such as slam, the API answers 400 with caption_hangul_text_latin_style for Korean text, because those styles have no Hangul glyphs.
Rules from the Video captions docs (read 2026-10-06).
What does each field do?
style picks the look and the motion, design changes that look for one request, font picks the face for a Hangul style, and language is only a speech hint. If you omit language, speech-to-text finds the language automatically.
| Field | Controls | Default when omitted |
|---|---|---|
style | Look and motion | slam for Latin text, black-outline for Korean text |
font | Hangul face for a Hangul style | The style's own face |
design | Per-request overrides of the style | The style's tokens |
language | Speech-to-text hint | Automatic detection |
Which styles take Korean?
Use black-outline, weight-shift, highlight, pill-karaoke, clip-wipe, editorial-emphasis or the ad look korean-ad. slam, punch and tiktok-green are Latin display faces and are rejected for Hangul text. A font value that is a Hangul face with a Latin style returns caption_font_requires_hangul_style.
Why is rejection the better outcome?
The docs note that a clip full of tofu boxes would cost the same as a good one, so Sume returns an error before the render and does not switch your style.
Sources
More in Media tools
- Check a Short has sound and picture before adding it to a YouTube show
YouTube asks shows for high-quality video and sound and says no-visuals videos likely will not count. Probe each Short with Sume video inspect first.
- compose_video_has_no_audio: a silent clip under a still on Sume
Timeline compose warns compose_video_has_no_audio and still completes. The output is silent. How to add a voice-over afterward with a Timeline render.
- Copy space in an AI image: leave room for text, overlay it in code
Ask for empty space in the prompt, check the region with Pillow, and put the headline on top in code, so the text is exact and the picture stays a picture.
- Custom thumbnail for each Short in a YouTube series: pull the frame
YouTube lets a Shorts series carry a unique custom thumbnail per Short. Pull the best frame from each render with Sume video frames, then dress it.
Written by Sume