Korean avatar captions: slam style returns 400, use a Hangul style
Sume rejects a Korean script with slam, punch or tiktok-green captions (400 caption_hangul_text_latin_style). Pick a Hangul style such as korean-ad.
Short answer
If you send a Korean script to Sume Avatar 1.0 with inline captions in the slam style, which is the default, the request fails with 400 caption_hangul_text_latin_style. Sume does not change the style for you. Those font faces render Hangul as tofu, the empty boxes you see when a font lacks a glyph. For Korean speech, choose a Hangul style such as korean-ad, weight-shift, black-outline, highlight, pill-karaoke, clip-wipe or editorial-emphasis.
Which styles work for Korean
The docs list the caption styles. Three Latin-style faces are rejected for Korean text; the others are the Hangul-ready ones.
| Style | Korean script allowed |
|---|---|
| slam (default) | No: 400 caption_hangul_text_latin_style |
| punch | No: 400 caption_hangul_text_latin_style |
| tiktok-green | No: 400 caption_hangul_text_latin_style |
| korean-ad | Yes: Hangul karaoke |
| weight-shift | Yes |
| black-outline | Yes |
| highlight | Yes |
| pill-karaoke | Yes |
| clip-wipe | Yes |
| editorial-emphasis | Yes |
The request that works
Send the captions object with a Hangul style on the talking-video request. Keep the language hint as auto unless you know it. Everything else in the request stays the same.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: kr-avatar-001" \
-d '{
"avatar_handle": "product_host",
"aspect_ratio": "9:16",
"quality": "plus",
"script": "안녕하세요. 오늘 새 제품을 소개합니다.",
"captions": {
"enabled": true,
"style": "korean-ad",
"language": "auto"
}
}'What else to know about inline captions
Inline captions burn into the clean final MP4 after generation and use the spoken script or video_inputs text. They do not create a separate billed caption job. A failure in the caption stage is a soft failure: the avatar job can still succeed with a clean video_url and captions.status set to failed. Sume never adds captions to preview stills. If you stored captions at preview create, they apply at generate-video time. The estimated duration for inline captions must also be 60 seconds or less.
- Choose the style first, then write the script.
- Check captions.status in the result, not only the job status.
- Use standalone Video captions for a video you already have.
Why this matters for disclosure
A disclosure line in Korean must be legible. If the caption style turns it into empty boxes, the line is useless. The 400 error protects you from shipping an unreadable disclosure. Pick the Hangul style in your template once, and every Korean clip gets readable text. A 20-second Korean clip at plus costs 20 x $0.245 = $4.90, and captions add nothing extra.
A Korean script has a practical length limit too. The planned duration must fit the 4 to 60 second window, and inline captions reject an estimated duration above 60 seconds. Spoken Korean runs at a different pace from English, so do not count words from an English script; time the Korean text by reading it aloud, or give each scene an explicit duration in video_inputs.
A short test before a batch
Send one 8-second Korean clip at standard first. It costs 8 x $0.184 = $1.47. Check the result for captions.status and look at the burned-in text at normal size on a phone. If the Hangul reads cleanly, reuse the same captions object for the whole batch. If you need a different look, switch to another Hangul style and run the same test again.
Sources
Related posts
More in Sume Avatar 1.0
- Lip sync a Korean voice recording on Sume: which route, what it costs
Avatar Video is English only in code. For a Korean recording, Sume offers audio-driven routes: H3 Max lip-sync and VEED Fabric. Windows, rates and a test plan.
- Live AI avatar vs recorded clip: who reviews the words first?
A live avatar speaks in real time; a Sume Avatar 1.0 clip is scripted, previewable and fixed. Why that matters while Tavus says disclosure features are coming.
- Longest Avatar 1.0 video is 60 seconds: $11.04 on standard
A 60-second Avatar 1.0 talking video costs $11.04 on standard, $14.70 on plus and $33.00 on max. The cap, the product-image rate and a split plan.
- Make an AI avatar from a prompt, traits or a photo: $0.95 a call
Sume Avatar 1.0 creates a reusable avatar from a text prompt, structured traits or a reference image. All three cost the same flat $0.95. See the bodies.
Written by Sume