Does Sume Avatar 1.0 speak Spanish or Korean? English only in code

The Avatar 1.0 talking-video prompt on main says English only. What that means for Spanish or Korean lines, and the audio-driven route to test instead.

5 min readSume
All posts

No for script-driven talking video: the Avatar 1.0 clip prompt on Sume's main branch tells the video model that the spoken language is English only and that non-English speech is not allowed. The public docs do not publish a language list for Generate avatar video, so treat English as the supported language and test anything else on a short sample before you plan a campaign around it.

What the repo and docs say about language, as of 2026-10-08
SurfaceWhat it saysConsequence
Talking-video clip prompt (code on main)Spoken language: English only; no non-English speechWrite the script in English
Talking-video docsNo language list publishedDo not assume Spanish or Korean works
Inline captionsTake a language hint; Hangul styles exist for Korean speechCaptions can be Korean even where speech is English
Sume TTS request schemaLanguage field; omitted defaults to English at the providerSet it for every non-English transcript
Lip-sync routesMouth follows the audio you supplyPossible path for non-English audio, test first

What the English-only line covers

The restriction sits in the prompt Sume builds for each spoken chunk of a script. That prompt asks for the exact English dialogue, forbids translating or improvising, and ends with a ban on non-English speech. It is a statement about what the pipeline asks the video model to do, not a measured accuracy figure, and it is not in the docs pages. Silence beats (voice.type: "silence") carry no speech, so they are unaffected.

Because the same pipeline plans clips at 2.8 words per second and 4 to 12 seconds per clip, the length checks are also tuned to English word counts. A Spanish script that is longer per second of speech can land outside the 4 to 60 second window sooner than the word count suggests.

Captions are the part that handles other languages

Inline captions are a separate stage. They run after generation, burn onto the clean MP4, and accept a language hint (default auto). The documented Hangul styles, such as korean-ad, exist for Korean speech, and Sume rejects a Korean script that asks for a Latin-only style such as slam with 400 caption_hangul_text_latin_style.

That makes a useful split. An English avatar clip can carry Spanish or Korean subtitles only if you author the cue text yourself through standalone Video captions with cues or segments, since inline captions follow the spoken script.

The audio-first route to test for a non-English line

If the speech must be Spanish or Korean, separate the voice from the picture. Generate the voice with Sume TTS and set its language field, then send that audio to a lip-sync route. The OpenAPI descriptions say the MiniMax H3 Max lip-sync route takes a still image or ready avatar plus Sume-hosted audio of 5 to 14.8 seconds and that the mouth follows the audio. They do not say which languages it handles well.

Price a trial before committing. The minimum billable clip on that route is 5 seconds at 480p.

  • 5 s x $0.0625 per audio second at 480p = $0.3125 for the smallest test.
  • 5 s x $0.10 at 768p = $0.50 for the same clip at the default tier.
  • Judge the result by watching the mouth against the audio, not by the success status of the job.

Decision rule

Use Avatar 1.0 talking-video for English scripts where you want Sume to plan scenes and stills. Use TTS plus a lip-sync route when the language is not English, accept that you are testing an undocumented combination, and keep the first run to one short line. Keep a record of the sample you approved so the same check applies to the next language.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume