Voice Design from a text description: Gemini vs ElevenLabs
Gemini TTS has Voice Design and 200 custom voices per project; ElevenLabs lists Voice Design on v4. Sume's docs cover cloning, not voice design.

Both vendors let you describe a voice in words and get one back. Google documents Voice Design from descriptions plus Voice Replication from reference audio in Gemini TTS, with storage rules; ElevenLabs lists Voice Design among Eleven v4 features and Instant Voice Clones from 10 seconds of audio. Sume's docs describe cloning an existing voice and creating avatars from a prompt, but I found no feature that designs a speaking voice from a text description.
The facts on each page
| Item | Gemini TTS | ElevenLabs Eleven v4 |
|---|---|---|
| Voice from a description | Voice Design | Voice Design (listed on the v4 page) |
| Voice from audio | Voice Replication from reference audio | Instant Voice Clone from 10 seconds; Professional Voice Clone supported |
| Stock voices | 30 prebuilt voices plus an extended library | 17,500+ voices |
| Custom voice storage | Stateful custom voices: max 200 per project, one-year retention. Stateless voice keys expire after 7 days | Not stated on the pages I read |
| Language | More than 130 (Flash TTS) | 90+ |
Writing a voice brief
A description works best when it reads like a casting note. Neither vendor page gives a template, so this is general practice, not a vendor rule:
- Age range, pace, and register: 'warm, mid-30s, unhurried, conversational'.
- The use: 'for a 20-second product ad', 'for a calm onboarding narration'.
- Language and accent, stated plainly.
- What to avoid: 'not announcer-style, no sing-song'.
- Do not ask for a named real person's voice. Describe qualities, not people.
Persistence matters more than it looks
Google's two storage modes change how you plan a campaign. A stateful voice persists for a year, up to 200 per project; a stateless key lapses after 7 days, so a recurring series that depends on the same persona must use the stateful mode or re-create the voice. Pin the voice id in your project notes with the date you made it.
In Sume, a voice you clone is stored in your workspace with a primary language, and TTS jobs reference it by voice.id. Cloning is not billed, but speaking is: Sonic is $0.0475 per 1K characters after list x1.25, before the 5.5% platform fee and cent rounding. If you need designed voices, make them at Google or ElevenLabs; if you need a consistent, cloned brand voice inside a Sume workflow, clone it there.
Sources
Related posts
More in Comparisons
- GitHub Actions schedule or a Sume schedule for a weekly video run?
For a weekly video run, GitHub Actions cron is UTC-first and best-effort at busy times; a Sume schedule has a timezone and a spend cap. You can also chain them.
- GPT-6.1 Sol structured outputs or a Sume Format output_schema?
GPT-6.1 Sol supports structured outputs and function calling in OpenAI's API. A Sume Format run adds media, a spend cap and a receipt. Pick by the job.
- GPT Image 2.5 mask edit, then an Ideogram 4.5 text pass: one chain
Use GPT Image 2.5 for the masked region change and Ideogram 4.5 for the text pass, both on POST /v1/images. When it beats one model, and what each step bills.
- GPT Image 2.5 medium is 6x cheaper than Nano Banana 2 1K on Sume
On Sume, GPT Image 2.5 at medium costs $0.0165 for 1024x1024 and Nano Banana 2 1K costs $0.10, a 6.1x gap. Ratios at 2K and 4K, and what the price omits.
Written by Sume