One voice in 23 languages: a test matrix to run before you commit
MAI-Voice-2.1 claims one voice with a native accent in 23 languages. Build a 23-line audition matrix and cost it before you promise it to a market.

Before you tell a market that one voice speaks its language, listen to 23 short lines, one per language, in a fixed order, with a score sheet. Microsoft says MAI-Voice-2.1 lets a single voice use all 23 languages with a native accent (read 2026-10-03). That is a claim about the model; whether it holds for your brand voice, your product names and your pronunciation list is something only a listening test can answer, and the test costs cents.
This post gives the matrix, the script to generate the request bodies, and the checks that catch the usual failures.
What the vendor claims
The MAI-Voice-2.1 page lists 23 languages: English (US, Australia, UK, India), Spanish (Spain, Mexico), Portuguese (Brazil, Portugal), French, German, Italian, Dutch, Danish, Finnish, Swedish, Norwegian (Bokmal), Chinese (Simplified), Korean, Turkish, Thai, Vietnamese, Russian, Hindi, Romanian, Hungarian, Czech, Indonesian and Polish. It also lists instant voice matching from short reference clips with no fine-tuning (model page, read 2026-10-03). The launch post adds that one voice keeps a native accent across all supported languages.
The audition sheet
Use one neutral line everywhere: a greeting, a price, a date and the product name. Have a native speaker write the localized line, not a translation tool, and keep the product name in its normal local form.
| Check | How to judge it | Fail when |
|---|---|---|
| Accent | Native listener rates 1 to 5 | Below 4 for a headline market |
| Identity | Same speaker across languages | Voice sounds like a different person |
| Names and numbers | Brand name, price, date, phone number | Product name read as a common word |
| Pace | Words per minute near the English read | Much faster or slower than the others |
| Artifacts | Clicks, breaths cut off, flat endings | Any in the first two sentences |
Generate the 23 requests
The script builds one request body per language for Sume's TTS Router and prints the character count and the cost at Sume's $47.50 per 1M. Sume's language field is a code such as ko or en, and the docs say regional tags compare by primary language. Fill the dict with your localized lines; two are shown.
import json
LANGS = ["en","es","pt","fr","de","it","nl","da","fi","sv","no","zh","ko","tr",
"th","vi","ru","hi","ro","hu","cs","id","pl"]
lines = {
"en": "Hello, this is the new travel mug. It costs 29 dollars.",
"es": "Hola, esta es la nueva taza de viaje. Cuesta 29 dolares.",
# fill the other languages with lines written by a native speaker
}
total_chars = 0
bodies = []
for code in LANGS:
text = lines.get(code)
if not text:
continue
total_chars += len(text)
bodies.append({
"model": "sonic-3.6",
"transcript": text,
"voice": {"id": "YOUR_VOICE_UUID"},
"language": code,
})
print(len(bodies), "of", len(LANGS), "languages filled")
print("characters:", total_chars)
print("cost at $47.50 per 1M: $%.4f" % (total_chars * 47.50 / 1_000_000))
print(json.dumps(bodies[0], indent=2))
Reading the results
A full audition of 23 lines of about 55 characters is roughly 1,300 characters, which is about six cents on Sume's rate and about two cents at Microsoft's Flash rate. The cost is not the constraint; the constraint is finding native listeners. Ask two per language for headline markets and one for the rest, give them the sheet above and a blind order, and do not tell them which engine made which clip.
Set the pass rule before you listen. A reasonable one: every headline market scores at least 4 on accent and identity, no language fails the names-and-numbers check, and no more than three languages score below 3 on anything. Languages that fail are not necessarily lost. Often the fix is in the text, not the engine: spell out numbers, write the brand name phonetically in that language, or shorten the sentence. Keep the winning text with the job id that produced it, so a later regeneration can be compared against the take you approved.
What Sume's language guard does for you
When a request names a language that differs from the voice's primary language, the API returns HTTP 409 tts_voice_language_mismatch before any job or charge, with voice_id, voice_language and request_language. In the agent tool it appears as a non-error warning that asks you to confirm and retry with confirm_language_mismatch: true. For an audition that is the right behavior: a voice whose library language is English being asked for Thai should stop and ask, and a deliberate cross-language test is the only time you confirm.
Whether a given Sonic voice handles a given language well is a listening question; the router docs do not promise it. Run the same matrix on every engine you are comparing and compare the score sheets, not the feature lists.
Sources
Related posts
More in Use cases
- Australia AI ad disclosure: no blanket rule, and the AANA review
Ad Standards says Australia has no blanket rule to disclose AI in ads. The AANA's code review asked whether to add one. What that means for a video ad today.
- Clipper channel with permission: does YouTube still count it?
YouTube's page says promoting others' content is not allowed even with permission. What a clipper channel must add, plus trim and caption mechanics.
- How many AI ad videos can you queue at once? Limits by plan
Sume accepts paid generation jobs until queue capacity runs out, then returns 429 queue_full. Static capacity: 6 Free, 24 Pro, 48 Startup, 120 Scale.
- beehiiv Content Import skips video posts: rebuild with Sume
beehiiv's Content Import skips audio and video posts, and iframe embeds do not show in email. Rebuild plan: poster frame, audio detach, web link.
Written by Sume