German and Italian text to speech API: de and it on Sume TTS
German (de) and Italian (it) are in both Cartesia Sonic 3.6 and Sume's voice library. Send language de or it, reuse one voice, and watch the 409 language check.

German and Italian are the easy cases among Sume's TTS languages: both are on Cartesia's Sonic 3.6 list and both are in Sume's own voice-library language list, so you can create, save and reuse a voice in either. Send language: "de" or language: "it" with the voice id and the Sume job does the rest.
Sume's voice-library list has 16 languages: en, ko, ja, zh, es, fr, de, pt, it, hi, nl, pl, ru, sv, tr and tl. German and Italian are in it, with Spanish, French and Portuguese, which is why a Western European release is the one case where the voice and the language line up in the library.
What is covered where?
Sonic 2 also lists German, but not Italian: its 2025 snapshots stop at English, French, German, Spanish, Portuguese, Chinese, Japanese and Korean, and Cartesia says those stop working after October 20, 2026. An Italian voice on Sonic 2 was never possible on those snapshots, so an Italian pipeline is on a Sonic 3 generation already.
| Language | Code | sonic-3.6-2026-08-27 | sonic-3.5-2026-05-04 | sonic-3-2026-01-12 | In Sume's voice-library language list |
|---|---|---|---|---|---|
| German | de | Yes | Yes | Yes | Yes |
| Italian | it | Yes | Yes | Yes | Yes |
| French | fr | Yes | Yes | Yes | Yes |
| Spanish | es | Yes | Yes | Yes | Yes |
| Portuguese | pt | Yes | Yes | Yes | Yes |
What does the language check do when I reuse a voice?
When the voice is saved in Sume, its primary language is compared with language. A known mismatch, such as a German-tagged voice asked to speak Italian, returns HTTP 409 tts_voice_language_mismatch with voice_language and request_language in the body, before a job or charge. The hosted MCP tool presents it as a warning with confirmation_required: true; only after the user agrees do you retry the same request and idempotency key with confirm_language_mismatch: true.
Regional tags compare by primary language, so de-AT against a German voice raises nothing. That also means the check will not tell you that Swiss or Austrian German is the wrong regional variant; it only compares the base language.
How do I produce both languages in one file?
Cartesia describes the language as a per-request setting, and Sume's contract does too, so a bilingual script is two requests, one per language, each with its own voice and idempotency key. Ask for WAV, then join the takes with Sume's audio timeline, which joins up to 20 parts without re-synthesis and without gaps at the seams. MP3 can leave encoder padding at a seam, so WAV is the safer container for anything you will concatenate.
- Price: $0.0475 per 1,000 characters; a 2,500-character German page is $0.119.
- Limits: 20,000 characters and 1,200 seconds of audio per job.
- Speed 0.6 to 1.5, volume 0.5 to 2, emotion as a short free-text guide up to 64 characters.
Request example
Italian, with a saved or raw voice UUID in VOICE_ID:
import os
import uuid
import requests
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={
"x-api-key": os.environ["SUME_API_KEY"],
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"transcript": "Benvenuti al nostro aggiornamento settimanale.",
"language": "it",
"voice": {"mode": "id", "id": os.environ["VOICE_ID"]},
"mode": "async",
},
timeout=30,
)
r.raise_for_status()
print(r.json())Which voice should I reuse?
One saved voice per language, named so that the language is obvious. Sume's voice tool records a primary language when you save a voice, and that is what the 409 check reads, so a voice saved with the right language is the one that protects you later. Use the voices list to find the voice id again instead of copying UUIDs between notebooks, and keep the same id across a series so every episode sounds alike.
Sources
Related posts
More in Models
- Google Pics API? Edit one object with Nano Banana on Sume
Google's Pics announcement describes an app, not an API. The closest call on Sume: Nano Banana reference edits, or a GPT Image 2.5 mask for one region.
- Google video model dates: Omni GA, Veo shutdowns, one table
One dated table of Google's video model lifecycle read from its own pages on Oct 3, 2026: Omni 1.1 Flash GA, Veo 3.1 preview shutdowns, and what to call.
- GPT-5.1, o3, GPT-5.4 Nano shutdown dates and the Sume model field
OpenAI lists shutdowns for gpt-5.1, gpt-5.3-codex, gpt-5.4-nano (Apr 1, 2027) and o3 (Dec 11, 2026). A Sume Format run with those ids gets a 400.
- GPT-6.1 Sol cache write is $2.50 per 1M: when does it pay off?
GPT-6.1 Sol lists $2 input, $2.50 cache write and $0.10 cached input per 1M tokens. One reuse of a prefix already beats sending it fresh; the math.
Written by Sume