MAI-Voice-2.1 has 23 languages, 26 locales, 28 codes: which to quote
Microsoft says 23 languages and 26 locales; OpenRouter lists 28 codes and says 30+. Quote 23 languages, and use Python to turn the 28 codes into 23.

Quote 23 languages. Microsoft's launch post says 23 languages and 26 locales, the OpenRouter page for the Flash model lists 28 locale codes yet says "30+ languages", and the 28 codes collapse to exactly 23 languages.
The four numbers
Different pages count different things, so a headline number depends on the page you read.
| Source | Number | What it counts |
|---|---|---|
| Microsoft launch post | 23 languages, 26 locales | Languages and regional variants |
| Microsoft model page | 23 languages | Language list |
| OpenRouter Flash page, codes | 28 | Locale codes such as en-US, en-GB, en-IN, es-ES, es-MX |
| OpenRouter Flash page, text | 30+ languages, 140+ voices | Marketing text, not matched by the code list |
28 codes to 23 languages
Take the primary subtag of each code. English has four codes (en-AU, en-GB, en-IN, en-US), Spanish two (es-ES, es-MX), Portuguese two (pt-BR, pt-PT), and the other 20 codes are one each, which gives 28 codes and 23 languages.
codes = """cs-CZ da-DK de-DE en-AU en-GB en-IN en-US es-ES es-MX fi-FI
fr-FR hi-IN hu-HU id-ID it-IT ko-KR nb-NO nl-NL pl-PL pt-BR pt-PT
ro-RO ru-RU sv-SE th-TH tr-TR vi-VN zh-CN""".split()
langs = {c.split("-")[0] for c in codes}
print(len(codes), len(langs))Which to put on a slide
Use 23 languages and cite Microsoft's model page. Use 28 only when you mean locale codes on OpenRouter, and avoid 30+ because the same page's code list does not support it. When you compare with Sume, compare language lists, not locale counts: Sume's voice library has 16 language tags.
Limits
Counts change as pages are updated. The figures here were read on 2026-10-05, and I did not test any voice.
Sources
Related posts
More in Models
- MAI-Voice-2.1-Flash: 150ms for 45 seconds of audio, for batch TTS
Microsoft says MAI-Voice-2.1-Flash makes 45s of audio at 150ms end-to-end latency, at $15 per 1M characters. What that does and does not tell a batch TTS user.
- MAI-Voice-2.1-Flash at 45 ms: does a rendered avatar need fast TTS?
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms of inference. A rendered avatar clip does not benefit from it. Where the latency shows up in a Sume job.
- MAI-Voice-2.1 Hindi: six voices listed, and a Hindi request on Sume
Microsoft's MAI-Voice-2.1 page lists six hi-IN voices with different style sets. Compare that with a Hindi request on Sume TTS and what the language field does.
- Mandarin Chinese text to speech API: MAI-Voice-2.1 zh-CN vs Sume zh
MAI-Voice-2.1 lists Chinese (Simplified) as zh-CN; Sume tags zh and bills by character. Audition Mandarin ad copy for 1 cent and know what to check.
Written by Sume