Korean caption styles on Sume: black-outline, clip-wipe and the rest
For Korean speech on Sume, pick a Hangul caption style: black-outline is the safe default, clip-wipe reads best small. All six, and why slam shows tofu.

For Korean speech, use a Hangul caption style on Sume video captions. black-outline is the safe default and what you get if you send Korean text with no style. clip-wipe is the clearest at small phone sizes. slam, punch and tiktok-green are Latin display faces with no Hangul glyphs, so the API rejects Korean text sent to them.
The six identities
korean-ad is the ad karaoke look: weight-shift with an accent colour on the spoken word. Sending no style does not resolve to it. It resolves to black-outline, so choose korean-ad on purpose.
| Style | Look |
|---|---|
black-outline | White fill on a thick black outline, middle of the frame; the safe default |
weight-shift | Phrase cards; the spoken word gets the heavy weight |
highlight | An accent block moves in behind the spoken word |
pill-karaoke | Card in a dark pill; colour changes with the voice |
clip-wipe | Each word wipes in left to right; clearest at small phone sizes |
editorial-emphasis | Two-line left-aligned card; the last word drops to a second line at about twice the size |
Why Latin styles are refused
A Latin face drawn over Hangul renders empty boxes, often called tofu, and that video costs the same as a good one. So the API returns 400 caption_hangul_text_latin_style instead of switching style behind your back. The same rule applies to font: a Hangul face with a Latin style returns 400 caption_font_requires_hangul_style.
The default render adds no gold accent either. If you want a colour on the spoken word, set design.colors.active.
Picking by clip
weight-shift and korean-ad animate the wght axis, and only Pretendard has that axis. On a static face they keep colour and scale emphasis but lose the weight travel.
- Talking head or UGC voiceover: start with
black-outline, then testclip-wipeif your audience watches on small screens. - Ad with a brand feel:
korean-ad, orweight-shiftfor the same weight animation without the preset accent. - Quote or editorial clip:
editorial-emphasis, which gives the last word of each phrase its own line. - Product demo with dark footage:
pill-karaoke, because the dark pill sits behind the text.
One request
A caption job costs $0.20 for a clip up to 60 seconds. Language is only a speech-to-text hint. It never selects the style or the font.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-ko-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "clip-wipe",
"language": "ko"
}'Test three looks for one transcription
Send source_caption_id instead of video_url with a new style. Sume reuses the source video and the word timings it already has, so speech-to-text does not run twice. A restyle is still a render, so the price does not change.
Related posts
More in Media tools
- Which Sume media step caps your clip: 1800, 900 or 300 s
Trim takes a 1800 s source, video filter and video frames stop at 300 s, compose at 300 s, Timeline at 1800 s. One table of every length cap.
- Whoosh and riser transition sounds from a music prompt on Sume
Sume has no sound-effects endpoint. Prompt the music router for a riser or hit, cut it with Timeline audio split at $0.01, and lay it under a cut.
- Wine tasting video music: a 2-minute refined, quiet bed
Brief a refined 2-minute instrumental for a wine tasting video that stays under the host's voice: one generation and a two-minute render, $0.325 on Sume.
- Winter coat lookbook clip with a size and care plate over the footage
Put a still size-and-care plate over a coat clip with Timeline compose overlay: width_ratio, margin_ratio and position set the plate, $0.02 a clip on Sume.
Written by Sume