MAI-Transcribe-2: 60 languages, 17 more than 1.5, and the Sume hint
Microsoft lists 60 languages for MAI-Transcribe-2, 43 for MAI-Transcribe-1.5, so 17 new ones. Sume STT takes a language_code hint or auto-detects.

MAI-Transcribe-2 supports 60 languages, and 17 of them were not supported by MAI-Transcribe-1.5: Afrikaans, Azerbaijani, Bosnian, Persian, Filipino, Galician, Hebrew, Armenian, Icelandic, Kazakh, Latvian, Macedonian, Malay, Nepali, Swahili, Urdu and Cantonese. The Learn page's language table has a column for each model, and counting the ticks gives 43 for 1.5 and 60 for 2 (read 2026-10-08).
If your audio is in one of those 17, the older model is not an option for you; the 60-language claim is also in Microsoft's launch posts.
What the table says
By default the streaming model runs in multilingual mode with language auto-detection, so you do not name the language. The launch post for the streaming model adds that detection is continuous, meaning it can follow a speaker who switches language mid-stream. I did not find a Microsoft page that publishes per-language accuracy for the streaming model; the 5.2% average word error rate in the batch post is a FLEURS average across languages, not a per-language figure.
| Model | Languages | Notes |
|---|---|---|
| MAI-Transcribe-1.5 | 43 | Counted from the table |
| MAI-Transcribe-2 | 60 | Table and launch posts |
| Added in 2 | 17 | Afrikaans to Cantonese, listed above |
| Includes Japanese and Arabic | Yes, both models | MAI-Voice has neither |
On Sume
Sume's STT 1.0 takes an optional language_code (BCP-47 or a provider language hint, up to 16 characters, for example en or ko). Omit it for auto-detect. The result echoes language_code and a language_probability, so you can log what was detected.
Sume's public docs do not publish a language list for STT 1.0 that I could check against Microsoft's 60, so test a 30-second sample in each language you serve before you commit. A one-cent job is the cheapest test there is.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"audio_url": "https://example.com/sample-sw.m4a",
"language_code": "sw", "duration_seconds": 30, "mode": "async"}'A test plan
Record or collect 30 seconds per language, run each as a job with and without a language_code, and compare the text. If the hinted run is clearly better, always send the hint. If neither is usable, that language is not covered for your purposes, regardless of any list.
What to do with a language outside the list
If a language is in neither list, do not assume a model will cope. Some models transcribe an unsupported language into the phonetically closest supported one, which looks like fluent but wrong text. Compare the output with a human transcript for a short clip, and check language_probability on Sume results: a low value on a clip you know is in one language is a warning.
For code-switching audio, where speakers mix two languages, test with real recordings. On Sume, send a single hint only if the clip is mostly one language, and otherwise omit it and review the result.
Sources
Related posts
More in Models
- MAI-Transcribe-2-Streaming runs in 4 regions, MAI-Voice in 14
Microsoft lists four regions for MAI-Transcribe-2-Streaming and 14 for MAI-Voice-2.1; three overlap. Sume's STT and TTS requests have no region field.
- MAI-Voice-2.1 languages: 23 listed, no Japanese or Arabic voice
Microsoft's MAI-Voice-2.1 voice table covers 23 languages and 28 locale codes, with no Japanese or Arabic row. How Sume's TTS language field handles both.
- MAI-Voice-2.1-Flash: 150 ms end-to-end and a 45-second audio limit
Microsoft's news post gives MAI-Voice-2.1-Flash 150 ms end-to-end latency and 45 s of audio; the Learn page gives no latency number. What each page supports.
- Nano Banana 2.1 vs Pro from 512 to 4K: price per image on Sume
Nano Banana 2.1 runs $0.075 at 512 to $0.20 at 4K; Nano Banana Pro is $0.1875 up to 2K and $0.375 at 4K. Full tier table and a per-1,000 view.
Written by Sume