Korean caption styles compared: weight-shift to editorial-emphasis

Sume has six Hangul caption identities plus korean-ad. Compare black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis.

5 min readSume
All posts

For Korean speech, Sume's caption job offers six Hangul identities: black-outline, weight-shift, highlight, pill-karaoke, clip-wipe and editorial-emphasis, plus korean-ad for the ad karaoke look. Start with black-outline unless you have a reason not to: it is what an omitted style resolves to for Korean copy, and the docs call it the safe default.

Everything below comes from Sume's video captions documentation. These are looks, not performance claims: the docs describe how each style behaves and make no statement about which one holds attention better, so test on your own audience.

How do the six identities differ?

korean-ad is separate: CapCut-style Hangul karaoke, one short phrase at a time in a lower third, with the spoken word shifting to a heavy weight. Pair it with language: "ko". It is not what an omitted style resolves to, so ask for it by name.

Sume Hangul caption identities as documented, read 2026-10-02
styleLook as documentedDefault face
black-outlineWhite fill on a thick black outline, mid-frameDo Hyeon
weight-shiftPhrase cards where the spoken word takes the weight and the rest drops backPretendard
highlightAn accent block sweeps in behind the word being spokenPretendard
pill-karaokeCard in a dark pill; colour tracks the voicePretendard
clip-wipeEach word wiped in left to right; clearest at small phone sizesDo Hyeon
editorial-emphasisLeft-aligned two-line card; the phrase-final word drops to a second line at about twice the sizePretendard lead, Black Han Sans emphasis

Which style suits which clip?

The docs give one explicit steer: clip-wipe is clearest at small phone sizes. Beyond that, the choice is yours, but the mechanics suggest some rules of thumb, which we label as such rather than as documented facts.

A talking-head or UGC clip with fast speech reads well in black-outline, because the thick outline keeps the text legible over busy backgrounds. A calmer voiceover where you want the eye to follow the voice fits pill-karaoke or highlight. editorial-emphasis is the most opinionated: it only makes sense where the last word of each phrase carries the point.

  • Weight travel in weight-shift and korean-ad animates the wght axis, which only Pretendard carries. On a static face those styles keep colour and scale emphasis but lose the weight travel.
  • editorial-emphasis always draws its emphasis line in Black Han Sans; font moves only the lead line.
  • An omitted style drops the gold spoken-word tint, so the spoken word keeps the fill colour. Naming black-outline keeps its own gold tint, and design.colors.active sets it either way.

How do you change one thing without switching style?

Every style is a set of design tokens, and the optional design object merges over them for one request. One key changes one thing, so a cyan emphasis on the safe default is a one-line override. Out-of-range numbers are a 400 at request time, so a bad look fails before it renders and bills.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ko-black-outline-cyan-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "black-outline",
    "language": "ko",
    "design": { "colors": { "active": "#22D3EE" } }
  }'

Can you try several styles without paying for transcription each time?

Yes. Pass source_caption_id instead of video_url with the new style, and Sume reuses the first caption's source video and its word timings, so no second speech-to-text runs. Billing is unchanged, because a restyle is still a render: you save the transcription step, not the $0.20 caption job price for videos up to 60 seconds under the current estimate. Pass words alongside it only to correct wording.

That makes a three-style comparison on one clip cheap in time even if it is not free in money. The first call transcribes; the next two restyle.

What errors should you expect?

Korean text on a Latin style (slam, punch, tiktok-green) returns 400 with caption_hangul_text_latin_style, because those faces have no Hangul glyphs. Naming a Hangul font with a Latin style returns caption_font_requires_hangul_style. design is not supported on punch or tiktok-green. A silent clip fails as caption_no_speech; pass cues instead. Sume does not substitute a style or font you did not choose, by design, so a failed request is the cue to fix the input rather than accept a surprise render.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume