One narrator in many languages: a Sume TTS batch vs MAI-Voice-2.1
MAI-Voice-2.1 claims one voice across 23 languages. On Sume, you send the same voice.id with a language per request and audition each language.

Microsoft's announcement says MAI-Voice-2.1 covers 23 languages and 26 locales with one voice usable across all of them, keeping a native accent. On Sume you get the same shape of workflow by sending the same voice.id with a different language on each POST /v1/tts-1.0/generate, one request per language, then listening to each result before you ship.
What the two sides promise
Microsoft states the language and locale counts and the one-voice claim on its announcement page. It is a vendor claim, and the voices page says the models are in public preview with no SLA.
Sume does not publish a language count for TTS. The request takes a language string of 2 to 16 characters (BCP-47 or ISO-639) and says to set it for every non-English transcript, because an omitted value defaults to English. A voice-language mismatch can raise a warning that you clear with confirm_language_mismatch: true; confirming does not change the voice or language.
| Item | MAI-Voice-2.1 | Sume TTS |
|---|---|---|
| Voice reuse | One voice, 23 languages, 26 locales | Same voice.id, language set per request |
| Language selection | By locale voice id | language field, 2-16 characters |
| Default if omitted | Depends on voice id | English |
| Mismatch handling | Not stated on the page | Warning, then confirm_language_mismatch |
| Status | Public preview, no SLA | Production endpoint |
A runnable batch skeleton
This builds one request body per language and prints them. Set VOICE_ID to a voice you have access to and swap the print for your HTTP client; it never reads or prints an API key.
import json, os
SCRIPTS = {
"en": "Welcome to the store.",
"es": "Bienvenido a la tienda.",
"ko": "\ud658\uc601\ud569\ub2c8\ub2e4.",
}
def bodies(voice_id):
for lang, text in SCRIPTS.items():
yield {"transcript": text,
"language": lang,
"voice": {"id": voice_id}}
voice = os.environ.get("VOICE_ID", "")
if not voice:
raise SystemExit("set VOICE_ID first")
for b in bodies(voice):
print(json.dumps(b, ensure_ascii=False))
When this does not fit
If the voice was not made for a language, an accent can leak through even when the request succeeds, and no setting on either side removes it. Listen to a sentence in each target language, especially for Korean, Japanese and tonal languages. Sume has no voice-cloning endpoint, so the voice has to come from an existing avatar or a library id.
For which endpoints produce the audio and how to poll them, see jobs and results.
A checklist for a multi-language batch
The failure modes are repetitive, so a small routine catches most of them.
- Keep one transcript per language in a file named by tag, so the
languagevalue comes from the filename and not from memory. - Run one short sentence per language first and listen before the full batch.
- Set
languageexplicitly for every non-English line; omitted means English. - For Korean, check the text really contains Hangul, or the API returns 400
tts_language_script_mismatch. - Keep the same
output_formatacross languages, so a later concatenation does not mix sample rates or encodings. - Name each output file with the language tag and the line number, which makes a bad take easy to find and redo.
Sources
Related posts
More in Use cases
- One Short in Korean and English: caption styles and rules
Posting a Korean and an English cut of your own AI Short? YouTube is silent on translations; Sume caption styles are not. Which style fits which script.
- One-take walkthrough prompts: store, kitchen, studio on Seedance 2.5
Three one-take walkthrough prompts for Seedance 2.5 (store aisle, restaurant kitchen, studio), each with a route, a pace and a 30 s price at 480p to 1080p.
- Open Enrollment Reminder Video With an AI Avatar: Script, Cost
Script an HR open-enrollment reminder as a Sume Avatar 1.0 clip: scenes, word and second counts, request JSON, and the cost on each quality tier.
- Open house invitation clip: MiniMax H3 Max, 8 seconds at 768p
A listing photo to an 8-second 16:9 open-house invitation with minimax-h3-max on Sume costs $0.80 at 768p, $0.50 at 480p and $1.60 at 1080p.
Written by Sume