Cartesia Sonic 3 languages: 44 in Sonic 3.6, how to set one
Cartesia says Sonic 3.6 speaks 44 languages, including new Odia and Urdu. How its language and locale fields work, and Sume's language field.

Cartesia's Sonic 3.6 text-to-speech model speaks 44 languages. Cartesia's docs list them by base code, from en, fr and ko to two new ones, Odia (or) and Urdu (ur). You choose one with the language or locale field, never both. Sume's TTS request has a language field only.
Cartesia's facts are from its Sonic 3.6 docs and launch post, read 2026-09-29. Sume's are from the TTS request schema in the Sume API reference.
Which 44 languages does Sonic 3.6 speak?
Cartesia's docs table lists these base language codes for the stable snapshot dated August 27, 2026: en, de, es, fr, ja, pt, zh, hi, ko, it, nl, pl, ru, sv, tr, tl, bg, ro, ar, cs, el, fi, hr, ms, sk, da, ta, uk, hu, no, vi, bn, th, he, ka, id, te, gu, kn, ml, mr, pa, or, ur.
Odia and Urdu are new in Sonic 3.6. Cartesia's launch post also says Hindi and English can be mixed in one generation (Hinglish), with the transcript written in Devanagari, Latin script, or a mix.
How do the language and locale fields work?
Per Cartesia's docs, language and locale accept a base code such as en or a regional code such as en-GB. Set one, never both. By default text normalization follows that value, so dates, times and numbers are read the local way; a separate normalization setting lets you use different conventions.
| Detail | What Cartesia says |
|---|---|
| Accepted values | A language code such as en or a regional locale code such as en-GB |
| Both fields at once | Set one, never both |
| Default normalization | Follows the value you set |
| Locale example | The optional locale field makes dates, times and numbers come out the way a local would say them |
How do I set the language on Sume?
Sume's TTS request has one field for this, language: a BCP-47 or ISO-639 code of 2 to 16 characters, such as ko, ja or en. The schema says to set it for every non-English transcript: omitted, it defaults to English at the provider, and Sume infers ko or ja only as a fallback from a Hangul or kana-only transcript.
The request schema lists no locale or normalization field and does not allow extra fields, so send only language. The Sume docs do not list which language codes each router model accepts; check the audio for any language you rely on.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sonic-36-es-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Su pedido sale mañana.",
"avatar_handle": "acme",
"language": "es"
}'What happens if the voice speaks another language?
The TTS schema has a confirm_language_mismatch boolean: set it to true only after the user confirms a voice-language mismatch warning. It does not change the voice or the language. For fixing wrong words rather than wrong languages, see Text to speech pronunciation.
Sources
Related posts
More in Models
- Text to speech API with emotion, speed and volume controls
Sume's text to speech API takes generation_config: speed 0.6 to 1.5, volume 0.5 to 2.0 and an optional emotion guide of up to 64 characters.
- Text to speech sample rate: which output_format to pick
Sume's TTS output_format takes mp3, wav or raw, six sample rates from 8000 to 48000 Hz and four PCM encodings. The default, and when to send wav.
- Veo 3 shutdown: which Gemini API models retired and what replaces them
Google retired veo-3.0-generate-001, veo-3.0-fast-generate-001 and veo-2.0-generate-001 on 2026-06-30. Its listed replacements are the Veo 3.1 preview models.
- Veo 3.1 4K API: 8 seconds only, and 4K options on Sume
Veo 3.1 makes 4K video only at 8 seconds, and not on the Lite model. On Sume, gemini-omni-flash-1.1 lists 4K for 3 to 10 seconds.
Written by Sume