Eleven v4 covers 90+ languages. Check your language on Sume TTS first
ElevenLabs lists 90+ languages for Eleven v4 and Cartesia lists 44 for Sonic 3.6, the engine under Sume TTS. How to test a language with the language field.

The counts
ElevenLabs' Eleven v4 post (read 2026-10-04) says the model supports 90+ languages. Cartesia's Sonic documentation (read 2026-10-04) lists 44 languages for Sonic 3.6, and Sume's POST /v1/tts-router/generate calls the Sonic models. So if your language is outside the 44, test before you plan around Sume TTS.
Counts do not guarantee quality. A language can be listed and still sound uneven on names, numbers or dialect, so a 20-second test beats a count.
The comparison
The numbers are from the vendors' own pages. The last row is the test, which only you can run.
| Service | Claim | What to do |
|---|---|---|
| Eleven v4 (ElevenLabs) | 90+ languages | Check your language and dialect in their voice library |
| Sonic 3.6 (Cartesia, used by Sume TTS router) | 44 languages, including Odia and Urdu | Send a real paragraph with the language field set |
| Sume TTS 1.0 | Language is an ISO code on the request | Set it for every non-English transcript |
How to test a language on Sume
Take a typical paragraph in your language, with a name, a number and a date in it. Choose a voice that speaks the language, and send the text with language set to its ISO code. If the voice and language disagree, Sume returns a mismatch warning, and confirm_language_mismatch lets you confirm on purpose; read the warning before you do, because it usually means the wrong voice.
Compare the result against a native speaker's reading. Listen for stress on names, how numbers are read, and whether punctuation produces natural pauses. Run the check on two voices, since one voice can fail where another is fine.
What to do when the language is not covered
Do not force an unsupported language through an English voice; the result is usually unusable. Use a vendor that lists your language, and note that moving the text between vendors is easy when your script is plain text and your voices are named in one config file. Sume's transcript field takes plain text, so a script written for one engine moves to another without conversion.
- Test at least 100 words, not a greeting.
- Try a second voice before you decide.
- Set
languageexplicitly rather than relying on guesses. - Re-read the vendor language page on the day you commit.
A short test script
Write the test text once in your language and keep it. Include a greeting, a price with a currency symbol, a date, a product name in Latin letters and a closing sentence with a question. Those five patterns catch most problems. Generate it on each candidate service at the same speed and listen back to back. The cost of the test on Sume is small: 400 characters is 0.4 x $0.0475 = $0.019.
Keep the file and the notes. When a vendor adds languages or ships a new model, repeat the same test and compare against your saved result.
Sources
Related posts
More in Comparisons
- Eleven v4 inline tags like [whispers] vs Sume's emotion guide field
Eleven v4 steers delivery with inline tags in the text. Sume TTS has a separate emotion string, plus speed and volume. How to port a tagged script.
- Eleven v4 Turbo 150 ms first speech vs Sume async TTS jobs
ElevenLabs quotes about 150 ms to first speech for v4 Turbo. Sume TTS is an async job, built for finished narration. Which one fits your use?
- ElevenLabs Music for a TV spot: two pages that do not agree
ElevenLabs music page says self-serve plans exclude film, TV and games; the API docs say cleared for film and TV. Get this in writing before a broadcast ad.
- ElevenLabs Music terms: restricted industries and banned inputs
Before an ad brief goes into ElevenLabs Music: the terms list restricted industries and prompt inputs you cannot send. What Sume does and does not enforce.
Written by Sume