One narrator in many languages: a Sume TTS batch vs MAI-Voice-2.1

MAI-Voice-2.1 claims one voice across 23 languages. On Sume, you send the same voice.id with a language per request and audition each language.

4 min readSume
All posts

Microsoft's announcement says MAI-Voice-2.1 covers 23 languages and 26 locales with one voice usable across all of them, keeping a native accent. On Sume you get the same shape of workflow by sending the same voice.id with a different language on each POST /v1/tts-1.0/generate, one request per language, then listening to each result before you ship.

What the two sides promise

Microsoft states the language and locale counts and the one-voice claim on its announcement page. It is a vendor claim, and the voices page says the models are in public preview with no SLA.

Sume does not publish a language count for TTS. The request takes a language string of 2 to 16 characters (BCP-47 or ISO-639) and says to set it for every non-English transcript, because an omitted value defaults to English. A voice-language mismatch can raise a warning that you clear with confirm_language_mismatch: true; confirming does not change the voice or language.

One voice, several languages (Microsoft announcement read 2026-10-05; Sume from the repo)
ItemMAI-Voice-2.1Sume TTS
Voice reuseOne voice, 23 languages, 26 localesSame voice.id, language set per request
Language selectionBy locale voice idlanguage field, 2-16 characters
Default if omittedDepends on voice idEnglish
Mismatch handlingNot stated on the pageWarning, then confirm_language_mismatch
StatusPublic preview, no SLAProduction endpoint

A runnable batch skeleton

This builds one request body per language and prints them. Set VOICE_ID to a voice you have access to and swap the print for your HTTP client; it never reads or prints an API key.

import json, os

SCRIPTS = {
    "en": "Welcome to the store.",
    "es": "Bienvenido a la tienda.",
    "ko": "\ud658\uc601\ud569\ub2c8\ub2e4.",
}

def bodies(voice_id):
    for lang, text in SCRIPTS.items():
        yield {"transcript": text,
               "language": lang,
               "voice": {"id": voice_id}}

voice = os.environ.get("VOICE_ID", "")
if not voice:
    raise SystemExit("set VOICE_ID first")
for b in bodies(voice):
    print(json.dumps(b, ensure_ascii=False))

When this does not fit

If the voice was not made for a language, an accent can leak through even when the request succeeds, and no setting on either side removes it. Listen to a sentence in each target language, especially for Korean, Japanese and tonal languages. Sume has no voice-cloning endpoint, so the voice has to come from an existing avatar or a library id.

For which endpoints produce the audio and how to poll them, see jobs and results.

A checklist for a multi-language batch

The failure modes are repetitive, so a small routine catches most of them.

  • Keep one transcript per language in a file named by tag, so the language value comes from the filename and not from memory.
  • Run one short sentence per language first and listen before the full batch.
  • Set language explicitly for every non-English line; omitted means English.
  • For Korean, check the text really contains Hangul, or the API returns 400 tts_language_script_mismatch.
  • Keep the same output_format across languages, so a later concatenation does not mix sample rates or encodings.
  • Name each output file with the language tag and the line number, which makes a bad take easy to find and redo.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume