MAI-Voice-2.1 languages: 23 listed, no Japanese or Arabic voice
Microsoft's MAI-Voice-2.1 voice table covers 23 languages and 28 locale codes, with no Japanese or Arabic row. How Sume's TTS language field handles both.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash speak 23 languages, and neither Japanese nor Arabic is one of them. Microsoft's news post says "23 languages and 26 locales"; the managed-voice table on the Learn page lists 28 locale codes across those 23 languages (counted on 2026-10-08). If you need ja or ar narration, MAI-Voice is not the model, and you need a different TTS source.
This post lists what the Microsoft pages show, then covers how Sume's text-to-speech route treats the language you ask for.
What Microsoft lists
The Learn page names the 23 languages through its voice IDs. They are Czech, Danish, German, English, Spanish, Finnish, French, Hindi, Hungarian, Indonesian, Italian, Korean, Norwegian Bokmal, Dutch, Polish, Portuguese, Romanian, Russian, Swedish, Thai, Turkish, Vietnamese and Chinese (Simplified).
English alone has four locales (en-AU, en-GB, en-IN, en-US), Spanish has two (es-ES, es-MX) and Portuguese has two (pt-BR, pt-PT). That is why locale codes outnumber languages. The news post counts 26 locales while the table I counted has 28; the table is the longer, newer list, and Microsoft's note on the page says it adds locales and voices as they become available.
| Item | Count | Where |
|---|---|---|
| Languages | 23 | News post and Learn page |
| Locales (news post) | 26 | microsoft.ai news post |
| Locale codes in the managed-voice table | 28 | Learn page, counted by hand |
| Japanese (ja-JP) rows | 0 | Learn page voice table |
| Arabic (ar) rows | 0 | Learn page voice table |
| Status | Public preview, no SLA | Learn page note |
Why it matters for a Japan or Gulf launch
MAI-Transcribe-2-Streaming lists Japanese and Arabic among its 60 speech-recognition languages, so the speech-in side of a voice agent covers both, while the speech-out side with MAI-Voice does not. A pipeline that transcribes Japanese callers with MAI and answers with MAI-Voice would need a second TTS vendor for the answer.
The Learn page also marks the models as public preview, offered without a service-level agreement and not recommended for production workloads. Plan a fallback voice either way.
How Sume treats the language
Sume's POST /v1/tts-1.0/generate has a language field, a BCP-47 or ISO-639 code. The API reference says to set it for every non-English transcript, because an omitted value defaults to English at the provider, and gives ko, ja and en as examples. Sume also infers ko or ja from a transcript that is only Hangul or kana, as a fallback.
If the chosen voice does not speak the language you requested, the request returns a mismatch warning that you must confirm with confirm_language_mismatch: true. Do not confirm it for production audio. The cheaper check is to audition the voice with a one-sentence job first; the 23-language audition matrix shows how.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"transcript": "こんにちは、新商品のご案内です。",
"language": "ja",
"voice": {"mode": "id", "id": "YOUR_VOICE_ID"},
"mode": "async"}'What Sume does not list
Sume does not list MAI-Voice models. Its TTS 1.0 route is billed at $0.0475 per 1,000 characters with a one-cent minimum per job and a 20,000-character cap per request. Arabic is covered in a separate post on the Arabic TTS API.
Decide on language first, price second: a cheaper per-character rate does not help if the voice is missing.
A checklist before you pick a voice vendor
Write down every language and locale you need, not just the language. Portuguese for Brazil and for Portugal are different rows on Microsoft's page, as are the English locales, and a voice listed for one locale is not a promise for its neighbor. Then check three more things: whether each locale has a voice in the style you need (several voices list only neutral), whether the vendor marks the model as a preview.
Run the same one-sentence test in every target language and have a native speaker listen. The count of languages on a page is a marketing number until someone who speaks the language has approved a take. On Sume this test costs 1 to 2 cents per line, so a 20-language check is well under a dollar.
Sources
Related posts
More in Models
- MAI-Voice-2.1-Flash: 150 ms end-to-end and a 45-second audio limit
Microsoft's news post gives MAI-Voice-2.1-Flash 150 ms end-to-end latency and 45 s of audio; the Learn page gives no latency number. What each page supports.
- Nano Banana 2.1 vs Pro from 512 to 4K: price per image on Sume
Nano Banana 2.1 runs $0.075 at 512 to $0.20 at 4K; Nano Banana Pro is $0.1875 up to 2K and $0.375 at 4K. Full tier table and a per-1,000 view.
- Same Omni video twice: no seed, but a replay returns it
Omni Flash lists no temperature or seed, and Sume rejects seed on every video model. To repeat a clip, replay an idempotency key or use a video_url edit.
- Seedance 2.5 has a 4-second minimum: the shortest clip is $1.0747
Seedance 2.5 on Sume takes 4 to 30 seconds. What a 4-second clip costs at 480p, 720p and 1080p on each Seedance tier, and where shorter hooks go.
Written by Sume