Korean captions on product clips: use korean-ad, not slam

slam, punch and tiktok-green have no Hangul glyphs and return 400 on Korean copy. Use korean-ad with language ko for Korean product clips; $0.20 a job.

5 min readSume
All posts

For a product clip where someone speaks Korean, burn captions with style: "korean-ad" and language: "ko" on Sume's video captions endpoint. It is the CapCut-style Hangul karaoke look: one short phrase at a time in a lower third, with the spoken word shifting to a heavy weight. A standalone caption job is $0.20 for videos up to 60 seconds under the current fixed estimate.

If you send Korean copy to slam, punch or tiktok-green, Sume refuses it with 400 caption_hangul_text_latin_style instead of rendering boxes. That is deliberate: the docs say a video full of tofu boxes would cost the same as a good one.

Why does the default Latin style fail on Korean?

Those three styles draw in Latin display faces with no Hangul glyphs. Rather than re-style your job into something you did not pick, Sume rejects it at request time, before a render is billed. The same applies to the font field: the Hangul faces are only valid next to a Hangul style, and naming one with a Latin style returns 400 caption_font_requires_hangul_style.

If you omit style entirely, the wording decides. Korean copy resolves to black-outline, which the docs call the safe default, and Latin copy resolves to slam. Note that korean-ad is not what an omitted style resolves to, so you must ask for it by name when you want the ad treatment.

Which Hangul caption style fits which clip?

The Hangul identities group words into phrase cards and burn the transcript exactly as written, with no case folding. The table repeats the docs' description for each, plus korean-ad. We have not benchmarked them against each other, so pick by look and test two on a small batch before you commit a whole catalog.

Pretendard is the default face for korean-ad, weight-shift, highlight, pill-karaoke and editorial-emphasis; Do Hyeon is the default for black-outline and clip-wipe. Any font value outside the documented list is rejected rather than substituted.

Hangul caption identities for Korean speech (Sume docs, read 2026-10-02)
StyleLookDefault face
black-outlineWhite fill on a thick black outline, mid-frameDo Hyeon
korean-adAd karaoke: weight-shift plus an accent colour on the spoken wordPretendard
weight-shiftPhrase cards; the spoken word takes the weightPretendard
highlightAn accent block sweeps in behind the spoken wordPretendard
pill-karaokeCard in a dark pill; colour tracks the voicePretendard
clip-wipeEach word wiped in left to right; clearest at small phone sizesDo Hyeon
editorial-emphasisLeft-aligned two-line card; the phrase-final word drops to a second line at about twice the sizePretendard lead line, Black Han Sans emphasis

What does the request look like?

The source must be a fetchable public HTTPS video URL. Localhost, private-network, signed or private URLs are rejected, as are non-HTTPS ones. Send an Idempotency-Key, then poll the job.

You can pass script_text when you want the burned wording to match your approved script. Sume keeps the speech-to-text timings as the source of truth and aligns your text to them; a bad alignment fails as script_alignment_mismatch or script_alignment_failed, and the suggested next action is to simplify the script or omit it.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ko-caption-sku-1001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/ko-product-clip.mp4",
    "style": "korean-ad",
    "language": "ko",
    "design": { "colors": { "active": "#22D3EE" } }
  }'

Can you restyle without paying for a second transcript?

Yes. If you burned korean-ad and want to compare clip-wipe, pass source_caption_id instead of video_url. Sume reuses that caption's source video and the word timings it already has, so no second speech-to-text runs. Billing is unchanged: a restyle is still a render at $0.20, so the saving is time and a second transcription, not money.

The design block above recolours the spoken word without changing anything else. It is accepted on the Hangul styles and slam, but not on punch or tiktok-green. For clips that have no speech at all, none of these apply; use authored cues instead, which skip speech-to-text.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume