Korean sale captions: slam and punch reject Hangul, use black-outline
Sume video-captions refuses Hangul text on slam, punch and tiktok-green with caption_hangul_text_latin_style. Pick black-outline or korean-ad; $0.20 per job.

If you send Korean text to a Latin caption style on Sume, the API returns 400 caption_hangul_text_latin_style and burns nothing. The slam, punch and tiktok-green styles use Latin display faces with no Hangul glyphs, and the video captions docs say the alternative would be a video full of tofu boxes at the same price as a good video. For Korean speech or Korean overlay cards, use black-outline, another Hangul identity or korean-ad. A job is $0.20 for a clip up to 60 seconds.
If you omit style, Korean text resolves to black-outline, not the Latin default.
Which style for which job
The docs list six Hangul identities, and korean-ad as the ad karaoke look for Korean speech, used with language: "ko". The language field only tells speech-to-text what to expect; it never selects the style or the font.
| Style | Look |
|---|---|
| black-outline | White fill on thick black outline, mid-frame; safe default |
| weight-shift | Phrase cards; spoken word goes heavy |
| highlight | Accent block moves behind the spoken word |
| pill-karaoke | Dark pill card, color changes with the voice |
| clip-wipe | Word wipes in left to right; clearest at phone sizes |
| editorial-emphasis | Two-line left card; last word at about twice the size |
| korean-ad | Ad karaoke look; use with language ko |
The font rule
The font field takes a Hangul face and works only with Hangul styles. Send one of those faces on a Latin style and the API returns 400 caption_font_requires_hangul_style. Both rules are request-time checks, so a wrong pairing fails before you pay for a render.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: kr-bf26-sku9-v1" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "korean-ad",
"language": "ko"
}'A bilingual sale clip
If one clip carries Korean speech and you also want an English cue for a price, split the work. Burn the Korean speech with a Hangul style first. The authored cues path does not run speech-to-text and burns text exactly as written, which suits a short price or a legal line. Remember that you send only one of script_text, words, cues and segments per job.
Check the output on a phone. Phrase cards sit in the lower third for korean-ad, so a platform interface may overlap them; the placement design tokens move the line if needed, and a restyle is still a render at the same price.
- Korean on slam, punch or tiktok-green is a 400.
- Omit style and Korean falls to black-outline.
- language is a speech hint, not a style.
Test a style on one clip first
Before you caption a whole Korean catalog, run one clip in two styles and compare on a phone. Each job is $0.20, so a two-style test is $0.40. Look for line breaks that split a word badly, a card that sits where the platform's own buttons will be, and a speech clip where the highlighted word drifts from the voice.
If the look is close but not right, do not switch styles at random. Use a design override on the same style: it changes colors, typography, placement, phrasing and motion for that one request, and each field you omit keeps the style's value. A number outside its documented range is a 400, so you find the mistake at submit time. Note that punch and tiktok-green do not support design.
Keep your sale clips consistent across the campaign. Choose one style and one set of overrides, write them down, and use the same body shape for each SKU. A Black Friday feed looks sharper when every clip shares one caption identity, and it makes the batch easier to review.
Related posts
More in Media tools
- LinkedIn video thumbnail under 2 MB: pull the frame with Sume
LinkedIn video ad thumbnails are JPG or PNG up to 2 MB. Pull a still from your clip with Sume's video frames endpoint and check its size before upload.
- Lip-sync a 20-second line: H3 Max takes 14.8 s of audio, so split it
MiniMax H3 Max lip sync on Sume accepts 5 to 14.8 seconds of audio. How to cut a 20-second line into two takes, or use Fabric for a longer face-to-camera line.
- LTX-2.5 48 fps option: conform frame rate on Sume
LTX-2.5 offers 24/25 fps or 48/50 fps. Sume's video catalog has no fps field, but Timeline output.fps conforms a clip to 24, 25, 30 or 60.
- Lyria 3.5 output: MP3 or WAV at 44.1 kHz, and what Sume returns
Google lists MP3 by default or WAV, 44.1 kHz stereo, with a SynthID watermark. Sume's music docs say the artifact is typically audio/mpeg. Check it in code.
Written by Sume