Can Sume Avatar 1.0 speak Spanish? English only, and what to do
Sume Avatar 1.0 speaks English only. Here are its documented limits, plus where to look when your viewers need Spanish or another language for the same ad.
No. Sume Avatar 1.0 speaks English only, so a Spanish script will not give you a Spanish presenter. If you need Spanish, French or any other language, do not send that script to the avatar route and hope. Plan a different route for the spoken part, and keep Avatar 1.0 for your English audience.
What does Avatar 1.0 accept?
The limits below come from the Sume docs for POST /v1/avatar-1.0/talking-video, read on the date in the caption.
| Setting | Documented value |
|---|---|
| Spoken language | English only |
| Length | 4 to 60 seconds, by Sume's estimate of the script |
| Input | One of script or video_inputs, not both |
| Quality | standard, plus (default), max |
| Aspect ratio | 1:1, 3:4, 9:16 (default), 4:3, 16:9 |
| Resolution | 720p |
What can I do for a Spanish audience?
Treat the language as a separate decision from the picture. Two Sume routes are written up on this blog: a text-to-speech language field with a voice-mismatch warning, and video models that can speak dialogue in several languages. Both are linked below. Neither is Avatar 1.0, and each has its own limits, so read the post before you commit a batch.
Whatever route you choose, test one short clip with a native speaker before you queue forty. Accent, pacing and brand names are the usual failure points.
Can I still use the avatar for English versions?
Yes. Use video_inputs for a hook, a silent demo beat and a call to action in one video. A voice.type: "silence" scene needs a duration and no script. Preview first-frame stills with Avatar video previews before you pay for the full render.
Voice cloning is an app feature and is not available through the API, so an API job uses the voices and avatars you have ready.
{
"avatar_handle": "your_ready_avatar",
"aspect_ratio": "9:16",
"quality": "plus",
"video_inputs": [
{ "id": "hook", "voice": { "type": "text", "script": "Wait, this fits in a coat pocket?", "duration": 3 } },
{ "id": "demo", "voice": { "type": "silence", "duration": 4 } },
{ "id": "cta", "voice": { "type": "text", "script": "Order today and it ships free.", "duration": 4 } }
]
}What can I do for Spanish audiences today?
Plan for the fact that this model covers English. These options do not need a different Avatar language.
- Make an English presenter clip for English-speaking audiences.
- Use a different part of your workflow for Spanish, such as a recorded voice-over on footage you already own.
- Add Spanish captions in your editing tool.
- Keep the English and Spanish versions as separate assets, so that nobody ships the wrong one.
Related posts
More in Sume Avatar 1.0
- Face swap beta takes no aspect ratio: crop 16:9 to 9:16 first ($0.02)
Sume's face swap beta has no aspect-ratio field. Crop a 16:9 source to 9:16 first with video filter ($0.02), then swap: $0.02 + $3.675 reserved at plus.
- Eight avatars on Sume: $7.60 to create, $117.60 for 24 plus clips
Eight Sume Avatar 1.0 characters cost 8 x $0.95 = $7.60 to create; three 20-second plus videos each (24 x 20 s x $0.245) add $117.60, $125.20 in all.
- Face swap beta, 9-second clip: you reserve the 15-second price
Sume's Avatar Face Swap beta reserves the maximum 15-second source price whatever your clip length: $2.76 at standard, $3.675 at plus, $8.25 at max.
- Griffin-Lite's 0.27 to 0.59 s latency in 25 fps frames vs Sume clips
Tavus reports Griffin-Lite video latency of 0.43 s average (0.27 best, 0.59 worst). At 25 fps that is 7 to 15 frames. Sume renders recorded clips.
Written by Sume