One voice in 23 languages: Microsoft's claim vs Sume's language guard

Microsoft says MAI-Voice-2.1 uses one voice across 23 languages. Sume sets language per job and warns on a voice mismatch. What that means when you buy.

4 min readSume
All posts

Microsoft says one MAI-Voice-2.1 voice can speak all 23 of its languages, with native accents, across 26 locales (Microsoft AI, read 2026-10-04). Sume works differently. A Sume TTS job takes one voice and one language, and if the voice does not fit the language it stops with tts_voice_language_warning until you confirm (API reference). The two are answers to the same question: what happens when text and voice disagree.

What each side promises

Cross-language behavior, from the vendor pages and the Sume API contract (read 2026-10-04)
QuestionMicrosoft MAI-Voice-2.1Sume TTS 1.0
Languages23 languages, 26 localesAny BCP-47 or ISO-639 code in language; the guard is per voice
One voice for all?Yes, per the announcementNot claimed. Pick a voice that speaks the language
Mismatch handlingNot describedtts_voice_language_warning; retry with confirm_language_mismatch
Price$22 per million characters$0.0475 per 1,000 characters

What the guard is for

A mismatched voice and language can still produce fluent audio, with the wrong accent, which is easy to miss in review. Sume turns that case into an error. The right response is to listen to a 200-character sample, then confirm only if the accent is what you want, for example a character voice.

How to test a multilingual claim yourself

Do not take a language count at face value. Take the same 200-character line in each of your target languages and listen with someone who speaks it. At Sume's rate each sample costs under a cent: 200 characters is $0.0095. Ten target languages cost about $0.095 to audition.

Include a number and a date in each sample line, since those are where a voice most often stumbles.

Decision rule

If your catalog needs the same brand voice in every market, favor a model that makes the one-voice claim and test it per language. If your catalog has local presenters, favor a stack that tracks a voice per language and refuses silently wrong pairs. Sume's setup suits the second case, and it adds the join, detach and render steps to the same bill.

Sources

Related posts

More in Models

All Models posts

Written by Sume