One voice in 23 languages: Microsoft's claim vs Sume's language guard
Microsoft says MAI-Voice-2.1 uses one voice across 23 languages. Sume sets language per job and warns on a voice mismatch. What that means when you buy.

Microsoft says one MAI-Voice-2.1 voice can speak all 23 of its languages, with native accents, across 26 locales (Microsoft AI, read 2026-10-04). Sume works differently. A Sume TTS job takes one voice and one language, and if the voice does not fit the language it stops with tts_voice_language_warning until you confirm (API reference). The two are answers to the same question: what happens when text and voice disagree.
What each side promises
| Question | Microsoft MAI-Voice-2.1 | Sume TTS 1.0 |
|---|---|---|
| Languages | 23 languages, 26 locales | Any BCP-47 or ISO-639 code in language; the guard is per voice |
| One voice for all? | Yes, per the announcement | Not claimed. Pick a voice that speaks the language |
| Mismatch handling | Not described | tts_voice_language_warning; retry with confirm_language_mismatch |
| Price | $22 per million characters | $0.0475 per 1,000 characters |
What the guard is for
A mismatched voice and language can still produce fluent audio, with the wrong accent, which is easy to miss in review. Sume turns that case into an error. The right response is to listen to a 200-character sample, then confirm only if the accent is what you want, for example a character voice.
How to test a multilingual claim yourself
Do not take a language count at face value. Take the same 200-character line in each of your target languages and listen with someone who speaks it. At Sume's rate each sample costs under a cent: 200 characters is $0.0095. Ten target languages cost about $0.095 to audition.
Include a number and a date in each sample line, since those are where a voice most often stumbles.
Decision rule
If your catalog needs the same brand voice in every market, favor a model that makes the one-voice claim and test it per language. If your catalog has local presenters, favor a stack that tracks a voice per language and refuses silently wrong pairs. Sume's setup suits the second case, and it adds the join, detach and render steps to the same bill.
Sources
Related posts
More in Models
- Pixal3D multi-view: prepare the input images
Pixal3D added multi-view inference in September 2026 under an MIT license. Sume has no 3D, but reference-image edit can prepare input views.
- PixVerse V6 native audio and camera work: what to check in an API
PixVerse's blog lists V6 with camera work and native audio, plus R2 and a $439M Series C total. How to test those claims against any video API's catalog.
- PixVerse V6 adds native audio: which Sume video ids make sound
PixVerse V6 ships camera work and native audio. The Sume video docs name no PixVerse id, so here is how to find models that generate audio and read the flag.
- Qwen-Image-2.1 is non-commercial: hosted options
Qwen-Image-2.1 has open weights but a Qwen Research License that is non-commercial unless licensed. Check the license first, then the hosted catalog.
Written by Sume