Korean talking avatar captions: black-outline or korean-ad, not slam
Korean script on a Sume avatar video with slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style. Use black-outline or korean-ad.
If the script of your avatar video is Korean, pick a Hangul caption style. Sume rejects Korean text on slam, punch and tiktok-green with 400 caption_hangul_text_latin_style, because those faces have no Hangul glyphs and would draw empty boxes. Use black-outline, which is the default when you omit style for Korean text, or korean-ad for the karaoke look, and send language: "ko" (Video captions).
Inline captions on the talking-video call
Generate avatar video takes an optional captions object that burns styled captions into the finished MP4 from your spoken script. It takes the same four settings as standalone captions: style, an optional font, a language hint and script_text.
| style | Look |
|---|---|
| black-outline | White fill on a thick black outline, middle of the frame; the safe default |
| korean-ad | Ad karaoke: weight-shift with an accent color on the spoken word |
| clip-wipe | Each word wipes in left to right; clearest at small phone sizes |
| slam / punch / tiktok-green | Latin faces: rejected for Korean text |
The request
Send the captions block with the rest of the body:
captions.enabled: truecaptions.style: "korean-ad"or"black-outline"captions.language: "ko", a speech-to-text hint; it never selects the style or the font
If the caption stage fails
A failure in the caption stage is a soft failure. The avatar job can still succeed with a clean video_url and captions.status=failed, and you can run standalone Video captions on that URL afterwards.
Sources
Related posts
More in Sume Avatar 1.0
- Library story hour: an 18-second avatar announcement, priced
An 18-second avatar announcement for a library event costs about $3.31 on Standard, $4.41 on Plus or $9.90 on Max at Sume's listed per-second rates.
- Lip-sync audio under 5 seconds: join two TTS lines with audio concat
MiniMax H3 Max lip-sync on Sume wants 5 to 14.8 seconds of audio. Join two short TTS lines with Timeline audio concat for $0.01 instead of padding.
- One avatar handle, three ratios: the same face in 9:16, 1:1, 16:9
Keep one AI presenter across Shorts, feed and webinar slots: reuse a single avatar_handle and change only aspect_ratio on each Sume talking-video call.
- Pipecat avatar TTFB metrics vs timing a Sume avatar job
Pipecat's Tavus service reports TTFB from TTSStartedFrame and BotStartedSpeakingFrame. A Sume avatar job has no first byte to time: measure submit to completed.
Written by Sume