Gemini TTS voices: 150+ or 2,000+? Two sources, one test plan
Google's changelog says 150+ Gemini TTS voices; a tracker says 2,000+. Do not pick by count. A voice test plan, and how voices are selected on Sume.

The two sources do not agree. Google's Gemini API changelog (read 2026-10-02) says the GA of Gemini 3.8 Flash TTS and Flash-Lite TTS on 2026-09-22 includes voice design, voice replication and 150+ prebuilt and custom voices. Digital Applied's September tracker says 2,000+. The two pages do not show whether the second figure counts languages, variants or custom voices, so use the vendor's number and ignore the count entirely when choosing.
A voice count does not help you ship. What helps is a voice that fits your script, language and brand, tested on your own lines.
How should I choose a voice?
Test with your own sentences, not the vendor demo. Pick three candidates, read the same ten lines in each, and listen for the things that fail in production: numbers, product names, and the end of long sentences.
| Step | What to do | Why |
|---|---|---|
| Shortlist | Pick 3 voices for the language | A count of 150+ or 2,000+ is not a quality signal |
| Same lines | Read the same 10 lines in each | Makes the voices comparable |
| Stress lines | Include numbers, names, a long sentence | These break first |
| Lock | Reuse one voice id for the whole project | Keeps timbre consistent |
How are voices selected on Sume?
Sume's TTS request takes a voice through the avatar_id or avatar_handle of a workspace avatar whose voice status is ready, or through voice.id when you already hold a voice id. If both are sent they must match, or the request returns 400. Set language for every non-English transcript; if you omit it, the provider defaults to English.
What if the voice and language do not match?
Sume returns a mismatch warning. Resend with confirm_language_mismatch set to true only after a person has confirmed it; confirmation does not change the voice or the language.
What should I do about the count in a brief?
Cite the vendor, say read 2026-10-02, and note that a tracker reports a different number. Then spend the time on the test, not the count.
How many voices do I actually need?
Most projects need one narrator, perhaps two for a dialogue, and a short list of alternates in case the first choice fails on a name or a number. Beyond that, a larger catalog adds search time, not quality. Decide your criteria first: language, age range, pace, and how formal the voice should sound.
Write the criteria down and score each candidate against them. When a stakeholder prefers a different voice, you have a record of why each was chosen, and you can re-test with a new script in minutes.
Does voice replication change the plan?
Google's changelog lists voice replication as part of the GA. A third-party tracker reports it is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India. If your audience or your team sits in one of those places, plan a prebuilt or designed voice instead, and check Google's current page before you rely on either report.
Sources
Related posts
More in Models
- Gemini TTS limit: 16,384 output tokens at 25 per second is 10:55
Gemini 3.8 TTS is reported to cap output at 16,384 tokens with 25 audio tokens per second, about 10 minutes 55 seconds. Sume's TTS cap is 1,200 seconds.
- gpt-5.6-terra and gpt-5.6-luna on Sume: which still run, as what
Sume retired GPT 5.6 from every picker. Terra has no successor and runs as itself, Luna moves to GPT-6 Luna only where that is admitted, Sol moves to GPT-6 Sol.
- GPT-6.1 Sol effort on Sume: picker offers Low, Medium, High, plus Fast
OpenAI lists five effort values for GPT-6.1 Sol. Sume's picker exposes three, keeps effort off the model id, and prices Fast as an opt-in. What that means.
- GPT Image 2.5 edit drifts after a few turns: repeat the preserve list
OpenAI says to repeat the preserve list on each iteration to reduce drift. How to run a one-change-per-call edit chain on Sume, which keeps no chat memory.
Written by Sume