Does Sume Avatar 1.0 speak Spanish or Korean? English only in code
The Avatar 1.0 talking-video prompt on main says English only. What that means for Spanish or Korean lines, and the audio-driven route to test instead.
No for script-driven talking video: the Avatar 1.0 clip prompt on Sume's main branch tells the video model that the spoken language is English only and that non-English speech is not allowed. The public docs do not publish a language list for Generate avatar video, so treat English as the supported language and test anything else on a short sample before you plan a campaign around it.
| Surface | What it says | Consequence |
|---|---|---|
| Talking-video clip prompt (code on main) | Spoken language: English only; no non-English speech | Write the script in English |
| Talking-video docs | No language list published | Do not assume Spanish or Korean works |
| Inline captions | Take a language hint; Hangul styles exist for Korean speech | Captions can be Korean even where speech is English |
| Sume TTS request schema | Language field; omitted defaults to English at the provider | Set it for every non-English transcript |
| Lip-sync routes | Mouth follows the audio you supply | Possible path for non-English audio, test first |
What the English-only line covers
The restriction sits in the prompt Sume builds for each spoken chunk of a script. That prompt asks for the exact English dialogue, forbids translating or improvising, and ends with a ban on non-English speech. It is a statement about what the pipeline asks the video model to do, not a measured accuracy figure, and it is not in the docs pages. Silence beats (voice.type: "silence") carry no speech, so they are unaffected.
Because the same pipeline plans clips at 2.8 words per second and 4 to 12 seconds per clip, the length checks are also tuned to English word counts. A Spanish script that is longer per second of speech can land outside the 4 to 60 second window sooner than the word count suggests.
Captions are the part that handles other languages
Inline captions are a separate stage. They run after generation, burn onto the clean MP4, and accept a language hint (default auto). The documented Hangul styles, such as korean-ad, exist for Korean speech, and Sume rejects a Korean script that asks for a Latin-only style such as slam with 400 caption_hangul_text_latin_style.
That makes a useful split. An English avatar clip can carry Spanish or Korean subtitles only if you author the cue text yourself through standalone Video captions with cues or segments, since inline captions follow the spoken script.
The audio-first route to test for a non-English line
If the speech must be Spanish or Korean, separate the voice from the picture. Generate the voice with Sume TTS and set its language field, then send that audio to a lip-sync route. The OpenAPI descriptions say the MiniMax H3 Max lip-sync route takes a still image or ready avatar plus Sume-hosted audio of 5 to 14.8 seconds and that the mouth follows the audio. They do not say which languages it handles well.
Price a trial before committing. The minimum billable clip on that route is 5 seconds at 480p.
- 5 s x $0.0625 per audio second at 480p = $0.3125 for the smallest test.
- 5 s x $0.10 at 768p = $0.50 for the same clip at the default tier.
- Judge the result by watching the mouth against the audio, not by the success status of the job.
Decision rule
Use Avatar 1.0 talking-video for English scripts where you want Sume to plan scenes and stills. Use TTS plus a lip-sync route when the language is not English, accept that you are testing an undocumented combination, and keep the first run to one short line. Keep a record of the sample you approved so the same check applies to the next language.
Sources
Related posts
More in Sume Avatar 1.0
- Synthesia needs a live consent clip: what do Sume photo avatars need?
Synthesia says personal avatars need a live consent recording. Sume's photo avatar takes an HTTPS image URL, so your own release process must cover consent.
- Synthesia Pro yearly, $64/mo: break-even vs Sume is 69.6 minutes
Synthesia lists Pro yearly at $64 a month for about 720 minutes a year. Sume standard matches that $768 outlay at about 69.6 minutes of video; plus at 52.2.
- Synthesia Starter: 12 minutes for $29, or the same 12 minutes on Sume?
Synthesia lists Starter at $29 for about 12 video minutes a month. Twelve minutes of Sume Avatar 1.0 costs $132.48 at standard. Where the gap comes from.
- Synthesia Studio Avatar costs $1,000 a year: what $1,000 makes on Sume
Synthesia lists Studio Avatar creation at $1,000 a year on annual plans. On Sume a new avatar costs $0.95, so $1,000 covers 1,052 avatar creations.
Written by Sume