Korean avatar video captions: fix caption_hangul_text_latin_style
A Korean script with captions style slam, punch or tiktok-green is rejected with 400 caption_hangul_text_latin_style. Use a Hangul style instead.
If your Avatar Video request has Korean script text and inline captions.style set to slam, punch or tiktok-green, Sume rejects it with 400 caption_hangul_text_latin_style. Pick a Hangul style such as korean-ad and resubmit.
The rule is stated in the Generate avatar video guide and the Video captions page, both read 2026-10-02.
Why is the request rejected instead of restyled?
The guide says those three faces render Hangul as tofu, the empty boxes you see when a font lacks a glyph. Sume refuses the request rather than silently re-styling it, so the style you named is the style you get or an error you can read.
The avatar guide lists slam as the default inline style, while the standalone captions page picks a style from the wording. Because the two pages differ, always name a Hangul style for Korean speech rather than relying on a default.
Which styles work for Korean?
Inline captions take the same four knobs as the standalone captions endpoint: style, optional font, a language hint and script_text.
| Style | Korean script |
|---|---|
| slam, punch, tiktok-green | Rejected with 400 caption_hangul_text_latin_style |
| korean-ad | Works: Hangul karaoke for Korean speech |
| weight-shift, black-outline, highlight | Works: Hangul identities |
| pill-karaoke, clip-wipe, editorial-emphasis | Works: Hangul identities |
What is the corrected request?
Only the captions block changes. Keep the script as it was.
{
"avatar_handle": "sume_clawra",
"script": "안녕하세요, 오늘 신제품을 소개합니다.",
"quality": "plus",
"captions": {
"enabled": true,
"style": "korean-ad",
"language": "ko"
}
}The standalone captions page pairs korean-ad with language: "ko", which is why the example sets it. For inline captions the guide's own sample uses language: "auto", so either is accepted; ko just removes a guess.
What if captions fail after the video renders?
That is a different failure. The caption stage soft-fails: the avatar job can still succeed with a clean primary video_url and captions.status=failed. You keep the video and can caption it later with a standalone job on the public video URL.
Remember two limits. Inline captions are rejected when the estimated duration is above 60 seconds, and they do not create a separate billed video-caption job. For design overrides, such as changing a highlight color, use the standalone endpoint; the inline block has no design field.
Checklist
Before you send a Korean avatar video, check these three things.
- Name a Hangul style; do not rely on a default.
- Keep the script at or under the 60 second estimate if captions are on.
- Treat
captions.statusas separate from the job status.
How do I handle a mixed Korean and English script?
The guide's rule is about Korean text, and it does not describe a threshold for mixed scripts. If your script contains Hangul at all, play it safe and name a Hangul style. The Hangul styles are built for Korean speech, and the page does not claim they are poor at embedded Latin words, but test a short render before a batch.
If you cannot be sure which language a user will type, branch in your code: detect Hangul characters in the script, and set style to a Hangul style when found, else use your Latin style. That keeps the 400 from ever reaching a user.
Can I preview the captions first?
Previews store caption intent on create and apply captions only at generate-video time; preview stills are never captioned. That means a preview cannot show you the caption look, but a bad style is still worth catching at the preview step if your code sets the captions block there.
For a look you can judge before the full render, caption an existing public video with the standalone endpoint. It costs $0.20 for videos up to 60 seconds under the current fixed estimate, and you can confirm live pricing in GET /v1/catalog and the OpenAPI.
Sources
Related posts
More in Sume Avatar 1.0
- Dub with lip sync: Meta Reels option vs Sume Avatar 1.0 (English-only)
Meta offers optional lip sync on translated Reels. Sume Avatar 1.0 is English-only, so a non-English talking shot uses TTS plus a lip-sync endpoint.
- Regenerate avatar preview stills, or start a new preview?
Regenerate refreshes first-frame stills from the stored preview request. Changing script, avatar, scene or aspect ratio needs a new preview. The full rule.
- One avatar handle, three platform cuts: keep the character consistent
Keep one presenter across LinkedIn, Snapchat and Pinterest by reusing a single avatar_handle and rendering each cut at the right ratio and length with Sume.
- Snapchat Spotlight needs 6 seconds: avatar clips that start at 4
Snap's Public Profile API takes Spotlight videos of 6 to 60 seconds, but Avatar 1.0 plans from 4. Plan scripts at 6 seconds or more and probe before posting.
Written by Sume