Gemini Flash TTS 130+ languages, Flash-Lite 100+: check yours first
Google lists 130+ languages for Gemini 3.8 Flash TTS and 100+ for Flash-Lite. A count is not a quality bar; test your language on Sume with 2 calls.

Google's speech-generation page lists 130+ languages for Gemini 3.8 Flash TTS and 100+ for Flash-Lite. Cartesia's Sonic 3.6 page lists 44. Those three numbers are not comparable as quality bars, since each vendor counts languages its own way and none of the pages say how natural each language sounds. If you need a specific language, the answer is a 1-cent listening test, not a headline count.
The counts, side by side
Sume routes Sonic, so the Sume side of the comparison is the Cartesia count. Sume's TTS language field takes a code, defaults to English when omitted, and compares against the primary language stored on the voice. A mismatch returns 409 tts_voice_language_mismatch before any job or charge, and you retry with confirm_language_mismatch: true only after you decide it is intentional.
| Model | Languages listed | Source page |
|---|---|---|
| Gemini 3.8 Flash TTS | 130+ | Google speech generation guide |
| Gemini 3.8 Flash-Lite TTS | 100+ | Google speech generation guide |
| Cartesia Sonic 3.6 (Sume router rows sonic-3.6, sonic-latest) | 44 languages | Cartesia models page |
A 1-cent test for any language
Write one 150-character sentence in your target language that includes a number, a name and a question. That is 150 x 0.00475 = 0.7125 cents, so the job costs 1 cent. Generate it with a voice whose stored language matches, listen once, and decide. Do this for the three languages you care most about, and you have spent 3 cents on a better signal than any count.
Pick sentences that expose real risk: dates, currency and a proper noun. A language that reads plain prose well can still stumble on a phone number.
import os, requests
r = requests.post("https://api.sume.com/v1/tts-router/generate",
headers={"x-api-key": os.environ["SUME_API_KEY"]},
json={"model": "sonic-3.6", "language": "es",
"transcript": "El pedido 4821 llega el 12 de marzo por 39,90 euros. Gracias, Lucía.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"mode": "sync", "wait_timeout_seconds": 30}, timeout=60)
print(r.status_code, r.json()["data"]["job"]["model"])When a bigger count matters
If your language is not on the Sonic list, a wider catalog matters more than any quality argument. In that case the router is not the right tool today: the Sume router seeds Cartesia Sonic only, and the docs list other TTS families as later catalog rows. Check the catalog with GET /v1/tts-router/models before you plan around a language.
Reading the result
Listen for four things: the numbers are read in the right language, the proper noun is not anglicised, the question has a rising or natural contour, and the currency word comes out right. Mark each as pass or fail and keep the audio. A language that fails two of four is not ready for customer-facing audio, whatever count the vendor page shows.
If you test several voices, repeat the same sentence so the only variable is the voice. Check the voice's stored language first: a voice made for English asked to speak Spanish returns the 409 warning, and that warning is a feature, not a bug, because it stops an accidental accent from being billed.
Budget for a full audition
Ten languages at 1 cent each is 10 cents. Even five voices across ten languages is 50 jobs, 50 cents. That is the cheapest research you will do on this decision.
Sources
Related posts
More in Comparisons
- Voice agent cost per minute: Gemini 3.8 Live vs gpt-realtime-2.1
Gemini 3.8 Live lists $0.005 per minute in and $0.018 out; gpt-realtime-2.1 lists $32 / $64 per 1M audio tokens. Convert tokens to minutes before you compare.
- Google's Omni-first video rule has an expiry: Veo 3.1 ends Oct 22
Google's Gemini API docs name Omni Flash the default for video and keep Veo 3.1 for extension, last-frame control and legacy pipelines until October 22.
- Omni 1.1 Flash 1080p vs Wan 3.0 1080p: price for 30 seconds
A 30 second 1080p film costs $7.50 on Sume's Wan 3.0 in one request, or $5.625 as three 10 s Gemini Omni 1.1 Flash clips. What the cheaper route costs you.
- Omni 1.1 Flash extends to 40 s; Seedance 2.5 makes 30 s in one pass
Google extends Omni 1.1 Flash clips in 10 s steps to 40 s. Seedance 2.5 generates 30 s in one pass. How to get 40 s on Sume with Seedance and Timeline.
Written by Sume