Captions on a non-English avatar clip: language hint and script_text
For Spanish, French or German scripts, set captions language to a hint or auto and pass script_text so the words match your script, not a mis-heard transcript.
Inline captions on a Sume avatar video take four knobs: style, an optional font, a language hint and script_text. For a Spanish, French or German script, leave language at auto or pass the matching code, and pass your own script text so the burned words match what you wrote. The language field only tells speech-to-text what to expect; it never picks the style or font. Latin-script languages use the default Latin styles without any extra setting.
Source: Generate avatar video and Video captions.
What the fields do
The docs spell out the split: style picks the look and motion, design edits it, font picks the face, and language is only a speech-to-text hint. With script_text present, captions align your words to the audio instead of trusting a transcript alone.
| Field | Use for a Spanish or French script |
|---|---|
style | Leave slam (default) or choose punch or tiktok-green; all draw Latin letters |
font | Hangul styles only; leave it out |
language | auto, or a hint code; it never changes the look |
script_text | Your exact script, so accents and brand terms are not mis-transcribed |
A request
The inline form sits beside the script in the same talking-video body, and it is not a separate billed caption job. The soft-fail rule matters: if the caption stage fails, the avatar job can still succeed with a clean video_url and captions.status set to failed.
{
"avatar_handle": "product_host",
"script": "Hola, esta es nuestra nueva botella térmica.",
"aspect_ratio": "9:16",
"captions": {
"enabled": true,
"style": "slam",
"language": "auto",
"script_text": "Hola, esta es nuestra nueva botella térmica."
}
}Check the result like a reader
Open the captioned MP4 with the sound off and read it. Accents, inverted punctuation and long compound words are where a wrong transcript shows up. If one language looks wrong, re-run the caption step on the clean video with the standalone route rather than regenerating the avatar.
- Estimated duration above 60 seconds is rejected for inline captions.
- Preview stills are never captioned.
- Keep the clean video; re-caption without a new render.
Limits to remember
The avatar docs list no supported-language set for avatar speech, so test your language with a short script first and judge the voice and lip movement yourself. Captions are the part the docs specify precisely; the voice is the part to verify by ear.
Sources
Related posts
More in Developers
- Backgrounded MCP tool lost progress: resume with Sume jobs_wait
A Claude Code fix covers MCP progress dropped when a tool moves to the background. Do not rely on progress for Sume jobs: re-issue jobs_wait in slices.
- Basin Pipelines: log Sume webhook deliveries and dedupe by job_id
Cloudflare Basin Pipelines streams ingest up to 1 GB/s. Log each signed Sume webhook delivery as one record and dedupe on job_id; Python sketch included.
- Blind-test Sonic 3.6 against 3.5 on your own script
A vendor's blind-test percentage is not yours. Render the same lines with two catalog versions through the TTS Router, shuffle them, and let listeners vote.
- Bluesky avatar and banner: 1,000,000 bytes, PNG or JPEG
Bluesky's profile record caps avatar and banner at 1,000,000 bytes each, PNG or JPEG only. How to size a GPT Image 2.5 output and check the bytes in Pillow.
Written by Sume