Sume Voices page lists 16 languages, Sonic 44: which list applies?
The Sume Voices page offers 16 languages when you create a voice; Cartesia states 44 for Sonic. The TTS request language field is a BCP-47 code you set per job.

The two lists answer different questions. The Voices page in the Sume web app offers 16 languages when you create a voice; the Sonic page from Cartesia states 44 languages for the speech model; and the TTS request just takes a BCP-47 language code that you set per job. Check all three if you create a voice and then synthesize in a language outside the 16.
The safest rule is to create the voice in the language you will use most, and audition it in every other language before you ship.
The three places a language shows up
In the Voices page, a language selector lists English, Korean, Japanese, Chinese, Spanish, French, German, Portuguese, Italian, Hindi, Dutch, Polish, Russian, Swedish, Turkish and Tagalog. In the TTS contract, language is a string of 2 to 16 characters; non-English requests must set it, and omitting it defaults to English at the provider. On the Cartesia Sonic page, the model is described as supporting 44 languages.
| Place | Count or form | What it controls |
|---|---|---|
| Sume web app, Voices page | 16 languages | Language chosen when creating a voice |
| Sume TTS request | BCP-47 code, 2 to 16 characters | Language the voice speaks the transcript in |
| Cartesia Sonic page | 44 languages | What the speech model covers |
What can go wrong
If you create a voice in English and ask it to read Korean, the request can return a mismatch warning. You can confirm it with confirm_language_mismatch, but the confirmation only acknowledges the warning; it does not change the voice or the language. Hearing the result is the only check.
- Create the voice in its main language.
- Set language on every non-English request.
- Audition each extra language with a native listener.
- Keep the model id fixed while testing.
A practical order
Decide the languages you need, check them against Cartesia's list, create or choose a voice in the closest of the 16 Voices languages, and then run a short audition matrix. If a language you need is outside the 16, test whether an existing voice speaks it acceptably rather than assuming it will not. The audition matrix post gives a layout.
Sources
Related posts
More in Models
- Tamil, Telugu, Kannada, Malayalam TTS API: language codes on Sume
Sonic 3.6 lists bn, ta, te, kn, ml, mr, gu and pa beyond Hindi. How to send each on Sume TTS, what the Hinglish note does not cover, and the cost per script.
- Sume image models at a $0.02 list price: Grok, Qwen, Imagen Fast
Three Sume image models list at $0.02 per image: Grok Imagine, Qwen Image and Imagen 4 Fast. How they differ on edits, ratios and image count per call.
- Veda sparse attention for MiniMax H3: 6.8x attention, 3.1x clip
Veda's sparse attention keeps 10 percent of attention work for MiniMax H3. Why 6.8x on attention becomes 3.1x per clip, and what it needs to run.
- Gemini 3.8 Flash is stable: keep model ids in config
When a vendor ships a new Flash model, a hard-coded id ages. Read Sume ids from GET /v1/catalog and treat 404 model_not_found as a signal, not a retry.
Written by Sume