Korean captions on product clips: use korean-ad, not slam
slam, punch and tiktok-green have no Hangul glyphs and return 400 on Korean copy. Use korean-ad with language ko for Korean product clips; $0.20 a job.

For a product clip where someone speaks Korean, burn captions with style: "korean-ad" and language: "ko" on Sume's video captions endpoint. It is the CapCut-style Hangul karaoke look: one short phrase at a time in a lower third, with the spoken word shifting to a heavy weight. A standalone caption job is $0.20 for videos up to 60 seconds under the current fixed estimate.
If you send Korean copy to slam, punch or tiktok-green, Sume refuses it with 400 caption_hangul_text_latin_style instead of rendering boxes. That is deliberate: the docs say a video full of tofu boxes would cost the same as a good one.
Why does the default Latin style fail on Korean?
Those three styles draw in Latin display faces with no Hangul glyphs. Rather than re-style your job into something you did not pick, Sume rejects it at request time, before a render is billed. The same applies to the font field: the Hangul faces are only valid next to a Hangul style, and naming one with a Latin style returns 400 caption_font_requires_hangul_style.
If you omit style entirely, the wording decides. Korean copy resolves to black-outline, which the docs call the safe default, and Latin copy resolves to slam. Note that korean-ad is not what an omitted style resolves to, so you must ask for it by name when you want the ad treatment.
Which Hangul caption style fits which clip?
The Hangul identities group words into phrase cards and burn the transcript exactly as written, with no case folding. The table repeats the docs' description for each, plus korean-ad. We have not benchmarked them against each other, so pick by look and test two on a small batch before you commit a whole catalog.
Pretendard is the default face for korean-ad, weight-shift, highlight, pill-karaoke and editorial-emphasis; Do Hyeon is the default for black-outline and clip-wipe. Any font value outside the documented list is rejected rather than substituted.
| Style | Look | Default face |
|---|---|---|
| black-outline | White fill on a thick black outline, mid-frame | Do Hyeon |
| korean-ad | Ad karaoke: weight-shift plus an accent colour on the spoken word | Pretendard |
| weight-shift | Phrase cards; the spoken word takes the weight | Pretendard |
| highlight | An accent block sweeps in behind the spoken word | Pretendard |
| pill-karaoke | Card in a dark pill; colour tracks the voice | Pretendard |
| clip-wipe | Each word wiped in left to right; clearest at small phone sizes | Do Hyeon |
| editorial-emphasis | Left-aligned two-line card; the phrase-final word drops to a second line at about twice the size | Pretendard lead line, Black Han Sans emphasis |
What does the request look like?
The source must be a fetchable public HTTPS video URL. Localhost, private-network, signed or private URLs are rejected, as are non-HTTPS ones. Send an Idempotency-Key, then poll the job.
You can pass script_text when you want the burned wording to match your approved script. Sume keeps the speech-to-text timings as the source of truth and aligns your text to them; a bad alignment fails as script_alignment_mismatch or script_alignment_failed, and the suggested next action is to simplify the script or omit it.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ko-caption-sku-1001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/ko-product-clip.mp4",
"style": "korean-ad",
"language": "ko",
"design": { "colors": { "active": "#22D3EE" } }
}'Can you restyle without paying for a second transcript?
Yes. If you burned korean-ad and want to compare clip-wipe, pass source_caption_id instead of video_url. Sume reuses that caption's source video and the word timings it already has, so no second speech-to-text runs. Billing is unchanged: a restyle is still a render at $0.20, so the saving is time and a second transcription, not money.
The design block above recolours the spoken word without changing anything else. It is accepted on the Hangul styles and slam, but not on punch or tiktok-green. For clips that have no speech at all, none of these apply; use authored cues instead, which skip speech-to-text.
Sources
Related posts
More in Use cases
- Find product moments in a live-commerce replay and trim them out
Run one transcript on a live-selling replay, find the sentences where each product is pitched, then cut each range with video-trim at $0.02 a clip.
- Live replay highlight reel in one render: detach once, no trims
Detach a live replay's audio once ($0.01), then build a 30-second highlight reel in one Timeline render using audio parts and source_in. No per-clip trim jobs.
- Log the exact video model id per shot after Runway model removals
When a vendor retires a video model, a logged model id turns the audit into a query. Record it on every shot, and know what Sume echoes.
- Meta AI Reels translation: who can use it, and the 1,000-follower rule
Per Meta, Reels translation is free for Facebook creators with 1,000 followers or more and for all public Instagram accounts where Meta AI is available.
Written by Sume