Pick an ElevenLabs model by language count: 90+, 70+, 32, 29
ElevenLabs lists 90+ languages for v4, 70+ for v3, 32 for Flash v2.5, 29 for Multilingual v2 and English only for Flash v2. Where Sume Sonic fits.

Match the model to the language list first: ElevenLabs says eleven_v4 covers 90+ languages, eleven_v3 70+, eleven_flash_v2_5 32, eleven_multilingual_v2 29, and eleven_flash_v2 is English only. None of those is a model id on Sume, whose TTS Router catalog is Cartesia Sonic only; Cartesia states 44 languages for Sonic.
If your script is in a less common language, the count decides before price or latency does.
The ElevenLabs list
The models doc gives the language counts below, along with latency and character limits where it states them. The text-to-speech capability page repeats Flash v2.5 at 32 languages with about 75 ms latency and Multilingual v2 at 29 languages, and lists MP3, PCM, mu-law, A-law and Opus as output formats.
| Model | Languages | Stated latency | Stated character limit |
|---|---|---|---|
| eleven_v4 | 90+ | Not stated | 10,000 |
| eleven_v4_turbo | 90+ | About 100 ms | Not stated |
| eleven_v3 | 70+ | Not stated | 5,000 |
| eleven_v3_conversational | 70+ | About 280 ms | Not stated |
| eleven_multilingual_v2 | 29 | Not stated | 10,000 |
| eleven_flash_v2_5 | 32 | About 75 ms | 40,000 |
| eleven_flash_v2 | English only | About 75 ms | 30,000 |
Where Sume fits
The Sume TTS Router offers sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. The Cartesia Sonic page states 44 languages, a number that sits between ElevenLabs' Flash v2.5 and v3 counts. The Sume request carries a BCP-47 language code, and the transcript limit is 20,000 characters, which exceeds the ElevenLabs limit for v4, v3 and Multilingual v2 but not for the Flash models.
Language counts are vendor claims about coverage, not quality. A language being listed does not tell you the voice sounds native, so audition it.
A selection routine
Write down the languages you must support at launch and the ones you might add. Remove any model that does not list the launch languages. Among what remains, run the same short script and compare. If the survivors are all ElevenLabs models, you will use ElevenLabs directly and import the audio into Sume for captions and timeline work. If Sonic covers your list, you can keep the whole flow on Sume.
- Check the exact locale you need, not just the language.
- Treat latency figures as streaming claims; Sume TTS is a job API.
- Keep the language code explicit on every Sume request.
- Re-check the lists in a month, since vendors update them often.
Sources
Related posts
More in Models
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
- Polish, Dutch, Swedish, Turkish text to speech API: Sume pl nl sv tr
Sume's Voices library tags voices pl, nl, sv and tr alongside 12 other languages. What to send for each, and how Eleven v4's list compares.
- QuantFunc INT4 MiniMax H3: 3.2 s per step on an RTX 4090
QuantFunc's 4-bit MiniMax H3 claims 3.2 s per step on an RTX 4090. That is not a clip time. What the card says, what it omits, and when to use a hosted job.
- Qwen Image Edit Plus: $0.03 a megapixel for text edits
fal prices Qwen Image Edit Plus at $0.03 per megapixel and highlights text editing and multi-image input. What that means for a Sume edit workflow.
Written by Sume