Korean captions API error: caption_hangul_text_latin_style (400)
Sume refuses Korean text on the slam, punch and tiktok-green caption styles with a 400. Use black-outline, korean-ad or another Hangul style; $0.20 a job.

If POST /v1/video-captions returns 400 caption_hangul_text_latin_style, you sent Korean text to a Latin caption style. slam, punch, and tiktok-green use display faces with no Hangul glyphs, so they would burn tofu boxes. Sume refuses the request instead of switching styles. Pick a Hangul style such as black-outline, korean-ad, or clip-wipe and resend.
Why the API refuses instead of falling back
The docs give the reason in one line: a caption style you did not select is worse than an error. The alternative is a video full of tofu boxes at the same price as a good video. The rule also covers font: a Hangul face sent with a Latin style returns caption_font_requires_hangul_style. Latin text on slam works as before.
Which styles to use for Korean speech
If you do not send style, the language of the caption text decides: Korean resolves to black-outline, Latin to slam. If you want the ad look, you must ask for it, since korean-ad is not the default. Karaoke-style accents are also not added to a default render.
| Style | Look |
|---|---|
| black-outline | White fill on a thick black outline, middle of frame; the safe default |
| weight-shift | Phrase cards; the spoken word gets the heavy weight |
| highlight | An accent block moves behind the spoken word |
| pill-karaoke | Dark pill card whose color changes with the voice |
| clip-wipe | Each word wipes in left to right; clearest at small phone sizes |
| editorial-emphasis | Two-line left-aligned card; last word enlarged in a display face |
| korean-ad | Ad karaoke look: weight-shift with an accent on the spoken word |
A working request
language is only a speech-to-text hint (ko, en) and never selects the style or the font. The video must be a public HTTPS URL that Sume can fetch.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ko-caption-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/ugc.mp4",
"style": "clip-wipe",
"font": "pretendard",
"language": "ko"
}'Choosing between Hangul styles
For talking-head and creator speech, black-outline is the safe default and needs no other fields. For ad creative, korean-ad gives the karaoke look with an accent on the spoken word. For small phone screens, the docs call clip-wipe the clearest. For a magazine look, editorial-emphasis puts the last word of each phrase on a second line at about twice the size in Black Han Sans.
Combine styles with design overrides to change a single token, for example the accent color, without changing anything else: design.colors.active. Numbers outside a documented range return a 400 at request time, so a bad look fails before you pay. Note that punch and tiktok-green do not support design.
Cost and fonts
Each accepted standalone caption job reserves and captures $0.20 for videos up to 60 seconds. A style or font error is a 400 at request time, so no job is accepted. Fonts are a fixed list (Pretendard, Do Hyeon, Black Han Sans, Jua, and others under the SIL Open Font License); any other name is rejected with no substitute. Note that weight-shift and korean-ad animate the weight axis, which only Pretendard has, so another face keeps the color and scale emphasis but loses the weight change.
Sources
Related posts
More in Media tools
- caption_no_speech: burn authored cues onto a silent clip, no ASR
A silent clip makes speech-to-captions fail with caption_no_speech. Send cues (text, start, end) and Sume burns your text at those times without speech-to-text.
- Check pause-cut seams with video frames: 24 stills cover 12 seams
Video frames returns up to 24 stills per call. Two stills per seam covers 12 seams per call, so a 14-cut Reel needs 2 calls to see every join.
- Concat 20 voice lines into one wav with timeline audio for 1 cent
Timeline audio concat joins up to 20 Sume-hosted audio parts into one gapless wav for a flat $0.01, and returns segment offsets to re-base your video slots.
- Conform a trim to 540x960 at 30 fps for TikTok non-Spark ads
TikTok's non-Spark ad spec sets 540x960 as the 9:16 minimum. One video-trim job with output 540x960 at 30 fps costs $0.02. Bitrate is not a field.
Written by Sume