Korean caption styles compared: weight-shift to editorial-emphasis
Sume has six Hangul caption identities plus korean-ad. Compare black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis.

For Korean speech, Sume's caption job offers six Hangul identities: black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis, plus korean-ad for the ad karaoke look. Start with black-outline unless you have a reason not to: it is what an omitted style resolves to for Korean copy, and the docs call it the safe default.
Everything below comes from Sume's video captions documentation. These are looks, not performance claims: the docs describe how each style behaves and make no statement about which one holds attention better, so test on your own audience.
How do the six identities differ?
korean-ad is separate: CapCut-style Hangul karaoke, one short phrase at a time in a lower third, with the spoken word shifting to a heavy weight. Pair it with language: "ko". It is not what an omitted style resolves to, so ask for it by name.
| style | Look as documented | Default face |
|---|---|---|
| black-outline | White fill on a thick black outline, mid-frame | Do Hyeon |
| weight-shift | Phrase cards where the spoken word takes the weight and the rest drops back | Pretendard |
| highlight | An accent block sweeps in behind the word being spoken | Pretendard |
| pill-karaoke | Card in a dark pill; colour tracks the voice | Pretendard |
| clip-wipe | Each word wiped in left to right; clearest at small phone sizes | Do Hyeon |
| editorial-emphasis | Left-aligned two-line card; the phrase-final word drops to a second line at about twice the size | Pretendard lead, Black Han Sans emphasis |
Which style suits which clip?
The docs give one explicit steer: clip-wipe is clearest at small phone sizes. Beyond that, the choice is yours, but the mechanics suggest some rules of thumb, which we label as such rather than as documented facts.
A talking-head or UGC clip with fast speech reads well in black-outline, because the thick outline keeps the text legible over busy backgrounds. A calmer voiceover where you want the eye to follow the voice fits pill-karaoke or highlight. editorial-emphasis is the most opinionated: it only makes sense where the last word of each phrase carries the point.
- Weight travel in
weight-shiftandkorean-adanimates thewghtaxis, which only Pretendard carries. On a static face those styles keep colour and scale emphasis but lose the weight travel. editorial-emphasisalways draws its emphasis line in Black Han Sans;fontmoves only the lead line.- An omitted
styledrops the gold spoken-word tint, so the spoken word keeps the fill colour. Namingblack-outlinekeeps its own gold tint, anddesign.colors.activesets it either way.
How do you change one thing without switching style?
Every style is a set of design tokens, and the optional design object merges over them for one request. One key changes one thing, so a cyan emphasis on the safe default is a one-line override. Out-of-range numbers are a 400 at request time, so a bad look fails before it renders and bills.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ko-black-outline-cyan-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "black-outline",
"language": "ko",
"design": { "colors": { "active": "#22D3EE" } }
}'Can you try several styles without paying for transcription each time?
Yes. Pass source_caption_id instead of video_url with the new style, and Sume reuses the first caption's source video and its word timings, so no second speech-to-text runs. Billing is unchanged, because a restyle is still a render: you save the transcription step, not the $0.20 caption job price for videos up to 60 seconds under the current estimate. Pass words alongside it only to correct wording.
That makes a three-style comparison on one clip cheap in time even if it is not free in money. The first call transcribes; the next two restyle.
What errors should you expect?
Korean text on a Latin style (slam, punch, tiktok-green) returns 400 with caption_hangul_text_latin_style, because those faces have no Hangul glyphs. Naming a Hangul font with a Latin style returns caption_font_requires_hangul_style. design is not supported on punch or tiktok-green. A silent clip fails as caption_no_speech; pass cues instead. Sume does not substitute a style or font you did not choose, by design, so a failed request is the cue to fix the input rather than accept a surprise render.
Sources
Related posts
More in Media tools
- Korean subtitle fonts for a video API: 29 faces on Sume
Sume's caption font field names 29 Hangul faces, all SIL OFL 1.1, from Pretendard to Gmarket Sans. Which suit which style, and what fails on a Latin style.
- Add a launch date plate to a teaser video with compose overlay
Put a launch-date card on the bottom of a teaser with Timeline compose overlay: width_ratio, margin_ratio, a 300-second ceiling and $0.02 a job.
- Launch teaser music on Lyria 3.5: section markers, 30 seconds
Write a 30-second launch-teaser track on Sume Music Router: section markers, one named moment and no duration field. $0.125 per generation.
- Loop background music under a long video with Timeline 1.0
A 30-second bed under a 90-second video: use soundtrack.loop, fade_out_seconds and duck_db in a Sume Timeline 1.0 render, and plan first for free.
Written by Sume