Gemini TTS: 30 prebuilt voices or 150+? Which number to quote
Google's docs say 30 prebuilt voices; its changelog says 150+ prebuilt and custom. Quote the one that fits and compare with Sume voice-id selection.

Two Google pages give two voice counts for Gemini 3.8 Flash TTS, and both are correct for different things. The speech-generation guide lists 30 prebuilt voices, an extended library, and a cap of 200 custom voices per project. The changelog entry for the Sep 22, 2026 general availability says 150+ prebuilt and custom voices. Quote 30 when you mean the voices you can name today from the guide, 150+ when you quote the changelog, and never add the two together.
What each page says
The guide names a few of the prebuilt voices (Zephyr, Puck, Charon and Kore) and mentions voice replication for custom voices. The changelog describes the total pool, which mixes prebuilt and custom. Neither page says the 200-per-project cap and the 150+ pool are the same quantity, so a safe sentence keeps them apart.
| Figure | Where | What it counts |
|---|---|---|
| 30 | Speech generation guide | Prebuilt voices |
| 150+ | Changelog, Sep 22, 2026 | Prebuilt and custom voices |
| 200 | Speech generation guide | Custom voices per project, maximum |
| 2 | Speech generation guide | Speakers in one request, maximum |
How Sume selects a voice
Sume does not publish a voice count, because a voice is something you own. A TTS request picks a voice with avatar_id, avatar_handle or voice.id, and a voice.id must be a TTS voice UUID or voi_ followed by 32 hex characters. Anything else returns 400 invalid_voice_id before a job is queued or a charge is made. The router does not create a second voice namespace, so a voice you made for TTS 1.0 works with model: "sonic-3.6" unchanged.
That makes the comparison a question of workflow, not catalog size. If you need to pick from a menu of named stock voices, a large prebuilt list helps. If you need the same voice for every video in a series, an id you control is the thing that matters.
A rule for quoting vendor counts
Write the figure, its source page and the date you read it, as the table caption does. If two pages disagree, quote both with their scope. A reader who checks the changelog next quarter will find a different number, and a dated caption tells them why.
Checklist before you publish a voice count
Open the page you are citing and confirm the exact sentence. Note whether the count includes custom voices. Note the date, because Google dated the general availability entry Sep 22, 2026 and a later edit can change the figure. Then state the scope in the same sentence as the number: 'Google's changelog says 150+ prebuilt and custom voices' is safe, while 'Gemini has 150 voices' is not.
The same discipline applies to Sume's own numbers. The 20,000-character job limit, the 1-cent minimum and the 95-cent maximum all come from the repository docs and the generated price catalog, so cite them by name when you write about them.
Sources
Related posts
More in Comparisons
- Two-voice dialogue: Gemini TTS vs ElevenLabs Text to Dialogue
Gemini TTS handles up to 2 speakers per request; ElevenLabs Text to Dialogue keeps continuity with surrounding text and request IDs. In Sume you stitch jobs.
- Voice Design from a text description: Gemini vs ElevenLabs
Gemini TTS has Voice Design and 200 custom voices per project; ElevenLabs lists Voice Design on v4. Sume's docs cover cloning, not voice design.
- GitHub Actions schedule or a Sume schedule for a weekly video run?
For a weekly video run, GitHub Actions cron is UTC-first and best-effort at busy times; a Sume schedule has a timezone and a spend cap. You can also chain them.
- GPT-6.1 Sol structured outputs or a Sume Format output_schema?
GPT-6.1 Sol supports structured outputs and function calling in OpenAI's API. A Sume Format run adds media, a spend cap and a receipt. Pick by the job.
Written by Sume