Voice one script in six languages: Sume TTS loop and 409 guard
MAI-Voice-2.1 sells one voice across 23 languages. On Sume, set language per line, handle the 409 voice-language guard and price a six-language batch.

To voice one script in six languages on Sume, translate the lines yourself, then submit one TTS job per language with the language field set to that line's language tag. If the chosen voice is not meant for that language, the API stops you with a 409 tts_voice_language_mismatch before it spends anything. Review the voice, or resend with confirm_language_mismatch: true if you accept the accent. The Python loop below does this for Spanish, French, German, Japanese, Korean and Portuguese.
The trend is Microsoft's MAI-Voice-2.1, launched 2026-10-01 with 23 languages across 26 locales and a single voice that speaks all of them, at $22 per 1M characters (read 2026-10-06 on the October tracker). Sume does not translate and does not claim one voice for every language. It speaks the text you send, in the voice you pick, and tells you when the two do not fit.
The loop
import json, os, time, urllib.error, urllib.request as u
KEY, HANDLE = os.environ["SUME_API_KEY"], os.environ["SUME_AVATAR_HANDLE"]
LINES = {"es": "Las ofertas empiezan el viernes.", "fr": "Les offres commencent vendredi.",
"de": "Die Angebote starten am Freitag.", "ja": "セールは金曜日に始まります。",
"ko": "할인은 금요일에 시작됩니다.", "pt": "As ofertas começam na sexta-feira."}
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
jobs = {}
for lang, text in LINES.items():
body = {"model": "sonic-3.6", "transcript": text, "avatar_handle": HANDLE, "language": lang}
try:
jobs[lang] = call("https://api.sume.com/v1/tts-router/generate", body, f"promo-v1-{lang}")
except urllib.error.HTTPError as e:
err = json.load(e).get("error", {})
if e.code == 409 and err.get("code") == "tts_voice_language_mismatch":
print(lang, "voice does not fit this language; pick another voice or confirm")
else:
raise
for lang, job in jobs.items():
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
arts = call(jobs[lang]["result_url"])["result"]["artifacts"]
print(lang, next(a["url"] for a in arts if a["type"] == "audio"))What the language field and guard do
Every non-English line needs the language field. If you leave it out, the provider defaults to English and the voice reads Korean or Spanish with English phonetics; the docs say never to rely on that. Sume infers ko or ja only as a fallback when the transcript is pure Hangul or kana.
Korean also has a script check. Romanized Korean sent with language: "ko" fails with tts_language_script_mismatch, because the text must be Hangul. Regional tags compare by primary language, so pt-BR and pt-PT are the same language for the guard; Filipino matches under fil or tl. A 409 is not an outage. It is a question: is this really the voice you want for this language?
Cost of a six-language batch
The price is per character, and the docs say it is rounded up to the cent. Each of the six lines above is about 30 characters, a tenth of a cent of raw price, so each job bills the one-cent floor and the batch is $0.06. For a 300-character line in six languages, that is 1,800 characters: $0.0855 raw on Sume's $0.0475 per 1,000, which becomes six jobs of $0.02 each, or $0.12 billed, (pricing, read 2026-10-06), against $0.0396 at MAI-Voice-2.1's tracker price of $22 per 1M and $0.027 at the Flash tier's $15 per 1M. Characters are counted as JavaScript string length, so a Japanese line counts per UTF-16 unit, and spaces and punctuation are billed.
Give each job its own Idempotency-Key. The script derives it from the language and a version label; derive it from the text as well if your copy changes between runs, or a rerun will return the old audio. Poll with next_poll_after_seconds, as in Jobs and results.
| Question | MAI-Voice-2.1 (tracker) | Sume TTS Router |
|---|---|---|
| Languages | 23 languages, 26 locales | Set language per request; the guard checks the voice |
| One voice across languages | Vendor claim | Voice-language guard returns 409 when it does not fit |
| Price per 1,000 characters | $0.022 (Flash $0.015) | $0.0475, same on every model |
| Translation | Not stated in the tracker | Not included; send translated text |
Decide once per voice and language
If you need one recognizable voice across markets, the guard is the feature to design around: test each language once with the voice you picked, listen, and only then set confirm_language_mismatch for the pairs you accept. Store that decision next to the voice id so a later batch does not ask again. See the guard in detail for the retry flow.
Sources
Related posts
More in Developers
- waitForJob on a failed Sume image job: read the record, skip /result
waitForJob resolves for failed and canceled jobs instead of throwing, and /result returns 409 job_not_completed for them. How to branch on job.status.
- Webhook for an unknown job_id: park it, then reconcile on insert
A Sume job.completed webhook can name a job_id your database has not stored yet. Return 2xx, park the event in SQLite, and apply it when the insert lands.
- Sume TTS source errors: which are safe to retry (status table)
Every tts_ error code in Sume's script-source API with its HTTP status, whether it charges, and whether a retry can help. A table for client error handling.
- Word timestamps to video frame numbers at 29.97 fps in Python
Sume STT words[] carry start and end in seconds. Convert them to frame indexes with exact 30000/1001 math so cuts do not drift on long timelines.
Written by Sume