One voice across 23 languages: MAI-Voice-2.1 vs a Sume avatar voice
MAI-Voice-2.1 keeps one voice across 23 languages. A Sume voice has one primary language and a 409 guard on mismatch. What that means for avatars.
MAI-Voice-2.1 is built so one voice speaks all 23 of its languages with a native accent. A Sume voice is built the other way round: it has one primary language, and the API stops a request that disagrees with it unless you confirm. For a presenter who must sound like the same person in several languages, the Microsoft model is closer to the goal on paper. For an avatar video where the wrong language would be an expensive mistake, the Sume guard is the safer default.
Two approaches to multilingual voice
| Item | MAI-Voice-2.1 | Sume (Cartesia Sonic) |
|---|---|---|
| Languages | 23 languages, 26 locales | Not checked for this post |
| Voice identity | One voice across all languages, native accent | Voice has a primary language, set at voices_create |
| Cloning | From just a few seconds of reference audio, with consent guardrails | Voice clone is free on Sume |
| Language mismatch | Not applicable | HTTP 409 tts_voice_language_mismatch, before any job or charge |
| Price | $22 per 1M characters | $47.50 per 1M characters (list x1.25), before fee |
How the Sume guard works
The target language comes from language on a REST request. If the stored voice language and the request differ, the API returns 409 with the voice_id, voice_language and request_language. You confirm by retrying with confirm_language_mismatch: true, without changing the voice, transcript or language. A matching request has no extra step, and regional tags compare by primary language, so a pt-BR request matches a pt voice.
The same selectors drive avatar work: TTS takes avatar_id or avatar_handle as its voice selector, and avatar videos are generated from the avatar handle. See the avatar docs for how an avatar is created.
Choosing for a multilingual presenter
- If one identical voice across many languages is the product, evaluate MAI-Voice-2.1 directly. Sume does not ship it.
- If you localise a video into a few languages, make one voice per language under the same avatar handle and let the 409 guard catch mismatched jobs.
- In both cases get a native listener to approve each language. Neither the language count nor a receipt proves pronunciation.
Sources
Related posts
More in Sume Avatar 1.0
- Pick a stock avatar by avoid_for and brand_safety_notes, not looks
Sume's avatar catalog returns profile metadata with best_for, avoid_for, brand_safety_notes and casting_notes. Read them before you cast a presenter for a clip.
- Griffin-Lite 26 of 54 Turing result: what the sample size says
Tavus reports 26 of 54 callers fooled by Griffin-Lite after a one-minute call. A 95% interval is about 35% to 61%. Python computes it, and says what to claim.
- Tavus Video to Face replica vs a Sume avatar from a photo or prompt
Tavus builds a replica from video or a photo; Sume builds an avatar from a prompt, props or a public photo URL. What each input gives you, and what it costs.
- What an AI avatar may not claim on TikTok Shop
TikTok Shop prohibits AI that impersonates real people or invents doctors and experts to endorse products. An avatar can present, but it can't be a fake expert.
Written by Sume