Describe a voice in words: Voice Design vs Sume Voices

ElevenLabs Voice Design takes 20 to 1000 characters. Sume Voices generate mode takes a 20 to 2000 character persona prompt and offers three variants.

4 min readSume
All posts

Both tools turn a written description into a new voice, with different limits: ElevenLabs Voice Design takes a description of 20 to 1000 characters, while the Sume Voices page in generate mode takes a persona prompt of 20 to 2000 characters and makes three variants for you to choose from. Sume's generate mode lives in the web app's Assets area, and the voice you save gets a library id you can reuse in TTS requests.

If you wanted a clone of a real person, that is a different feature; this one invents a voice from words.

What ElevenLabs documents

The ElevenLabs voices doc lists Voice Design with a description of 20 to 1000 characters. It also lists Instant Voice Clone from under two minutes of audio and Professional Voice Clone, which needs a Creator plan or above and can be shared in the Voice Library. The Eleven v4 post separately mentions clones from about 10 seconds, so check which number applies to your plan.

What Sume's Voices page does

In the Voices page of the web app, generate mode asks for a persona prompt and a gender (male, female or nonbinary), and a language chosen from a list of 16: English, Korean, Japanese, Chinese, Spanish, French, German, Portuguese, Italian, Hindi, Dutch, Polish, Russian, Swedish, Turkish and Tagalog. The prompt must be at least 20 and at most 2000 characters, and the page produces three variants to audition. A saved voice carries an id beginning voi_ followed by 32 hex characters, which the TTS contract accepts as voice.id.

Describe-a-voice limits side by side (read 2026-10-03)
ItemElevenLabs Voice DesignSume Voices, generate mode
Description length20 to 1000 characters20 to 2000 characters
OutputA voice to preview and saveThree variants to choose from
Where usedElevenLabs products and APISume TTS via voice.id
Clone from a real sampleSeparate featureSeparate mode on the same page

Writing a good description

Say who the speaker is, how they sound and how they should deliver lines. Age range, pace, warmth and setting beat vague adjectives. Avoid naming a real celebrity; you want an original voice that you have the right to use.

  • State the use: narration, ad read, tutorial.
  • Describe pace and energy in plain words.
  • Mention accent only if you have tested it in your target language.
  • Pick the language before the persona, since the voice is tied to it.

After you save

Run the same ten-line script through all three variants, and keep the winner's id with a note about the language you tested it in. For delivery tweaks like speed and emotion, use the generation_config fields in the TTS request rather than rewriting the persona.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume