MAI-Voice-2.1 languages vs Sume voice tags: 14 shared, 9 and 2 apart
Microsoft lists 23 TTS languages; Sume's voice library tags 16. Fourteen overlap, nine are MAI-only, two (Japanese, Tagalog) are Sume-only. The full split.

Fourteen languages appear on both lists. Microsoft's MAI-Voice-2.1 page lists 23 languages, and the voice library that Sume's mismatch guard reads has 16 language tags. Nine MAI languages have no Sume tag (Thai, Vietnamese, Indonesian, Danish, Finnish, Norwegian, Czech, Hungarian, Romanian). Two Sume tags are not on Microsoft's list (Japanese and Tagalog).
The three groups
The MAI side comes from the language list on Microsoft's model page. The Sume side comes from the VOICE_LANGUAGES constant that tags every voice in the Sume voice library: en, ko, ja, zh, es, fr, de, pt, it, hi, nl, pl, ru, sv, tr and tl.
| Group | Count | Languages (ISO 639-1 style) |
|---|---|---|
| On both lists | 14 | en, es, pt, fr, de, it, nl, zh, ko, hi, tr, ru, pl, sv |
| MAI-Voice-2.1 only | 9 | th, vi, id, da, fi, nb (Norwegian Bokmal), cs, hu, ro |
| Sume voice tags only | 2 | ja, tl |
| Union | 25 | everything above |
What a tag means on Sume
A tag is metadata on a voice, not a hard gate. Sume's language field on TTS takes any string from 2 to 16 characters. When the voice has a known primary language and the request names a different one, the API answers 409 tts_voice_language_mismatch before it creates a job or charges anything; you confirm and retry with confirm_language_mismatch: true. When the voice's metadata is unknown, a raw voice ID still goes through. Regional tags such as es-MX compare by primary language, and fil and tl count as the same language.
Sume's TTS Router (POST /v1/tts-router/generate) takes a required model (sonic-3.6, sonic-3.5, sonic-3, sonic-latest or sonic-preview), one of transcript or transcript_source, a voice selector and a language string of 2 to 16 characters. It is character-metered at $0.0475 per 1,000 characters, rounded up to whole cents per job, with a 20,000-character cap per request.
Reproduce the split yourself
The counts above come from plain set arithmetic, so you can re-run it when either list changes.
mai = "en es pt fr de it nl zh ko hi tr ru th ro hu cs da fi id pl sv nb vi".split()
sume = "en ko ja zh es fr de pt it hi nl pl ru sv tr tl".split()
print(len(mai), len(sume))
print("both", len(set(mai) & set(sume)))
print("mai only", sorted(set(mai) - set(sume)))
print("sume only", sorted(set(sume) - set(mai)))How to use the split
If your campaign sits inside the 14 shared languages, you can pick on price and workflow rather than coverage. If it needs one of the nine MAI-only languages, Microsoft's model is the documented option today, and a Sume job in that language is an experiment you should audition first (one 210-character line costs one cent). If it needs Japanese or Tagalog, Microsoft's page does not list them, while Sume's voice library does.
What this does not tell you
A tag says what a voice was labelled for, not how it sounds on your brand names or numbers. Neither list says anything about accent quality. Count languages after you listen to the audition, not before. Sume also returns a finished file per job (no streaming, no SSML field), while Microsoft's Flash model is built for live agents, so the two are not substitutes for the same job.
Sources
Related posts
More in Comparisons
- Flash TTS '55% faster inference' vs the end-to-end time of a TTS job
Microsoft quotes 55% faster inference and 150 ms end to end for MAI-Voice-2.1-Flash. Sume TTS is an async job; here is what you can actually time.
- MAI-Voice-2.1 at $22 per million characters vs Eleven v4
MAI-Voice-2.1 is $22 per 1M characters and Flash $15; Eleven v4 is $22 per 1M in a promo through Oct 12. Sume's Sonic route is $47.50 per 1M characters.
- MAI-Voice styles via express-as vs Sume's emotion field
MAI voices set styles with SSML mstts:express-as; some only have neutral. Sume has no SSML; generation_config.emotion is a free string up to 64 characters.
- Max video length in October 2026: TikTok, YouTube Shorts, Sume tools
TikTok's API allows 10 minutes and YouTube Shorts 3. Sume caps timeline and inspect at 1800 s, trim output at 900 s, and frames and filter at 300 s.
Written by Sume