MAI-Voice-2.1 lists 23 languages; how Sume TTS sets its language
Microsoft lists 23 languages and 26 locales for MAI-Voice-2.1. Sume TTS takes one language field, defaults to English and asks you to confirm a voice mismatch.

Microsoft documents MAI-Voice-2.1 as covering 23 languages and 26 locales (Microsoft AI, read 2026-10-09). Sume TTS has a single language field that defaults to English, and the OpenAPI reference expects you to set it for any other language. When the chosen voice does not match, Sume returns a language-mismatch response that you accept with confirm_language_mismatch. The two designs are different, and the language count alone does not settle which one you need.
What each side states
Microsoft's announcement gives the language and locale counts for MAI-Voice-2.1 and a per-character price. Sume's public pages do not give a language count for TTS in the pages I read, so this post makes no claim about Sume's total; check the voice you intend to use.
| Item | MAI-Voice-2.1 | Sume TTS 1.0 |
|---|---|---|
| Languages stated | 23 languages, 26 locales | No count stated in the pages read |
| Language control | Voice names include a locale, such as en-US-Harper:MAI-Voice-2.1 | language field; English by default |
| Mismatch behavior | Not described in the text read | confirm_language_mismatch to proceed |
| Price | $22 per 1M characters | $0.0475 per 1,000 characters |
Setting language correctly on Sume
For a non-English script, send the language field with the script. Forgetting it is the common mistake: the request defaults to English and the voice reads foreign text with English rules. For a voice whose own language differs from the script language, Sume stops with a mismatch response instead of producing odd speech, and you either pick a better voice or confirm on purpose.
Brand names are the other trap. An English brand name inside a Korean or Spanish script is read by the chosen language's rules, so check the audio. Pronunciation dictionaries, through pronunciation_dict_id, are the documented lever for terms that must be said a particular way.
A selection rule
Count the languages you actually ship, not the languages a vendor lists. Test one real paragraph per language with the voice you plan to use, and listen. If your list is inside Microsoft's 23, you can price that route at $22 per 1M characters. If your pipeline already runs on Sume jobs and the language works on the voice you picked, the job-based route saves a transfer step.
What to verify yourself
Before you commit, run the same paragraph in each target language through the voice you plan to use and have a native speaker listen. Check names, numbers and dates, which are where text-to-speech errs most. For Microsoft, confirm the locale you need is among the 26. For Sume, confirm the voice and language pair does not trigger the mismatch response. A one-paragraph test of 500 characters costs $0.02375 on Sume and $0.011 on MAI-Voice-2.1.
Record the result with the date, because both products are changing quickly: MAI-Voice is in public preview and its language list may grow.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 voice cloning is gated; how Sume TTS picks a voice
MAI-Voice-2.1 clones a voice from a 5 to 60 second clip but needs gated access. Sume TTS selects voices by avatar or voice id and has no reference-audio field.
- MAI-Voice-2.1 emotion control vs Sume TTS emotion, speed and volume
MAI-Voice-2.1 lists emotion control. Sume TTS takes generation_config: emotion (1-64 chars), speed 0.6-1.5, volume 0.5-2. A three-take test costs 3 cents.
- MAI-Voice-2.1 zero-shot voice prompting vs a Sume voice id or avatar
MAI-Voice-2.1 lists zero-shot voice prompting. Sume TTS takes a voice id or an avatar's voice, not a sample clip. What each request can and cannot express.
- MiniMax H3 768p vs Gemini Omni Flash 720p: an 8-second clip
On Sume, an 8-second clip is $0.60 on minimax-h3 at 768p and $1.00 on gemini-omni-flash-1.1 at 720p, with $0.30 at Omni's 360p. Price and constraints.
Written by Sume