Korean captions: slam returns 400, so use korean-ad with language ko
Latin caption styles have no Hangul glyphs, so Korean text on slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style. Use korean-ad.

Korean text sent to slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style. Sume refuses rather than switching the style, because the alternative is a video full of tofu boxes at the same price. For Korean speech use style: "korean-ad" with language: "ko", or one of the Hangul identities.
Hangul styles
From the video-captions docs, read 2026-10-09.
| Style | Look |
|---|---|
| black-outline | white fill, thick black outline, mid-frame; safe default |
| korean-ad | ad karaoke: weight shift plus accent on the spoken word |
| weight-shift | phrase cards, spoken word heavy |
| highlight | accent block behind the spoken word |
| pill-karaoke | dark pill, color follows the voice |
| clip-wipe | left-to-right wipe, clearest on small phones |
| editorial-emphasis | two-line card, last word large in a display face |
Defaults and language
If you send no style, the caption text decides: Korean resolves to black-outline, Latin to slam. language is only a speech-to-text hint (ko, en); it never selects the style or the font. korean-ad is not the default for Korean, so name it if you want the ad look.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ko-ad-001" \
-d '{"video_url":"https://media.sume.com/artifacts/example/talk.mp4","style":"korean-ad","language":"ko"}'Fonts
font is optional and only for Hangul styles; sending a Hangul face with a Latin style returns caption_font_requires_hangul_style. Pretendard is the face for korean-ad, weight-shift, highlight, pill-karaoke and editorial-emphasis; Do Hyeon is for black-outline and clip-wipe. The weight-travel styles animate the wght axis, which only Pretendard has.
Cost
The standalone caption job is $0.20 for clips up to 60 s under the current estimate. A wrong style fails at request time with a 400, so you do not pay for it.
Handling the 400 in a pipeline
Treat caption_hangul_text_latin_style as a routing signal. If your pipeline sends mixed-language copy, detect Hangul in the transcript or script, choose korean-ad or black-outline before the request, and keep slam for Latin text. The two default styles carry a gold tint on the spoken word when you select them by name; with no style the render keeps the fill color. design.colors.active sets the tint in either case. Every media job follows the same lifecycle: submit with an Idempotency-Key, receive a job, poll GET /v1/jobs/:id/status until it is ready, then read GET /v1/jobs/:id/result. A retry with the same key does not queue a second job, so a network error during submit never doubles a charge.
Sources
Related posts
More in Media tools
- LinkedIn 1200x675 video ad: Sume needs even sizes, so use 1280x720
LinkedIn lists 1200x675 as a 16:9 size, but 675 is odd. Sume timeline sizes must be even, so render 1280x720 or 1920x1080, both inside LinkedIn's range.
- LinkedIn 30-minute video: 24 spot-check stills, one every 75 seconds
LinkedIn allows videos up to 30 minutes. Video inspect takes 1,800 seconds and 24 stills per call, so at[] with 37.5 + 75k covers the whole file in one job.
- LinkedIn 30-minute video ad at 500 MB: 2.2 Mbps average, $3.00 render
A 30-minute LinkedIn ad must average under about 2.2 Mbps to fit 500 MB. Sume's timeline renders up to 1,800 seconds for $3.00; trim stops at 900.
- LinkedIn 9:16: 720x1280 is listed, Sume's 1080x1920 default is 2.25x
LinkedIn lists 720x1280 for 9:16 with a 1080x1920 max. Sume's Timeline default is 1080x1920, 2.25x the pixels; set output 720x1280 and fps 25.
Written by Sume