Voice Design from a text description: Gemini vs ElevenLabs

Gemini TTS has Voice Design and 200 custom voices per project; ElevenLabs lists Voice Design on v4. Sume's docs cover cloning, not voice design.

4 min readSume
All posts

Both vendors let you describe a voice in words and get one back. Google documents Voice Design from descriptions plus Voice Replication from reference audio in Gemini TTS, with storage rules; ElevenLabs lists Voice Design among Eleven v4 features and Instant Voice Clones from 10 seconds of audio. Sume's docs describe cloning an existing voice and creating avatars from a prompt, but I found no feature that designs a speaking voice from a text description.

The facts on each page

Vendor pages, read 2026-10-05
ItemGemini TTSElevenLabs Eleven v4
Voice from a descriptionVoice DesignVoice Design (listed on the v4 page)
Voice from audioVoice Replication from reference audioInstant Voice Clone from 10 seconds; Professional Voice Clone supported
Stock voices30 prebuilt voices plus an extended library17,500+ voices
Custom voice storageStateful custom voices: max 200 per project, one-year retention. Stateless voice keys expire after 7 daysNot stated on the pages I read
LanguageMore than 130 (Flash TTS)90+

Writing a voice brief

A description works best when it reads like a casting note. Neither vendor page gives a template, so this is general practice, not a vendor rule:

  • Age range, pace, and register: 'warm, mid-30s, unhurried, conversational'.
  • The use: 'for a 20-second product ad', 'for a calm onboarding narration'.
  • Language and accent, stated plainly.
  • What to avoid: 'not announcer-style, no sing-song'.
  • Do not ask for a named real person's voice. Describe qualities, not people.

Persistence matters more than it looks

Google's two storage modes change how you plan a campaign. A stateful voice persists for a year, up to 200 per project; a stateless key lapses after 7 days, so a recurring series that depends on the same persona must use the stateful mode or re-create the voice. Pin the voice id in your project notes with the date you made it.

In Sume, a voice you clone is stored in your workspace with a primary language, and TTS jobs reference it by voice.id. Cloning is not billed, but speaking is: Sonic is $0.0475 per 1K characters after list x1.25, before the 5.5% platform fee and cent rounding. If you need designed voices, make them at Google or ElevenLabs; if you need a consistent, cloned brand voice inside a Sume workflow, clone it there.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume