Cartesia voices speak up to 25 languages: voice plus language on Sume
Cartesia says 50+ library voices speak up to 25 languages natively. On Sume you send a voice id and a language code per job, and you test each pair.

Cartesia says more than 50 of its library voices speak up to 25 languages natively, checked against a native-speaker rubric. On Sume you pick a voice id and set a language per job, and the platform warns when the pair looks mismatched, so a multilingual voice means you test each language you need rather than trusting a count.
Treat the vendor's 25 as an upper bound for the best voices and as a prompt to audition.
What Cartesia says
The Cartesia multilingual voices post (2026-09-23) describes 50+ library voices speaking up to 25 languages natively, voice cloning from ten seconds of audio, and a native-speaker rubric used to judge them. The wording is "up to": not every voice covers all 25.
What the Sume request needs
The TTS contract takes a voice from avatar_id, avatar_handle or voice.id (a UUID or a voi_ library id), and a language as a BCP-47 code of 2 to 16 characters. Omitting the language defaults to English at the provider, with a fallback that infers Korean or Japanese from a Hangul-only or kana-only transcript. Never rely on that fallback for other languages; set the language every time.
When the voice and language look inconsistent, the request returns a warning and you resend with confirm_language_mismatch set to true after a person has agreed. The confirmation does not change the voice or the language you asked for.
| Item | Source | Meaning |
|---|---|---|
| 50+ voices | Cartesia multilingual voices post | Library voices speaking multiple languages |
| Up to 25 languages | Cartesia multilingual voices post | Best case per voice |
| language field | Sume TTS contract | BCP-47 code, 2 to 16 characters |
| confirm_language_mismatch | Sume TTS contract | Set after the user accepts the warning |
A matrix test
List your target languages in rows and your candidate voices in columns. Generate the same 15-second script, translated, for each cell and mark the ones you would ship. Do not extend a passing result to a neighbouring language; accent and pacing change between them. The 23-language audition post has a ready-made layout.
- Use a script with numbers and a proper name in every language.
- Have a native speaker listen, as Cartesia's own rubric does.
- Record voice id, language and model id from the job result.
- Re-test after any model change.
Cloned voices
Cartesia's cloning from ten seconds is a Cartesia product feature. Sume documents no public cloning endpoint, as the earlier voice cloning post explains, so plan around the voices you can select today.
Sources
Related posts
More in Models
- Cartesia Sonic 3.6: 44 languages or 61 locales, which to quote
Cartesia states 44 languages on its Sonic page and 61 locales in the Sonic 3.6 launch post. How to quote it, and what the Sume language field accepts.
- ChatGPT Try On from a screenshot: the same edit through an API
ChatGPT Try On starts from a selfie plus a product screenshot. Do the same edit with openai/gpt-image-2.5 on Sume: two references, one prompt, one Python call.
- Chinese, Hindi and Russian text to speech API: Sume zh, hi and ru
Sume's Voices library has zh, hi and ru tags. How to request each, what Eleven v4 lists, and the one field that stops an English-sounding read.
- claude-fable-5 is still active at Anthropic; Sume runs it as Fable 5.1
Anthropic keeps claude-fable-5 active until Jun 9, 2027, but a Sume Format run with that id runs on Fable 5.1. Vendor dates, price and the catalog alias.
Written by Sume