Describe a voice in words: Voice Design vs Sume Voices
ElevenLabs Voice Design takes 20 to 1000 characters. Sume Voices generate mode takes a 20 to 2000 character persona prompt and offers three variants.

Both tools turn a written description into a new voice, with different limits: ElevenLabs Voice Design takes a description of 20 to 1000 characters, while the Sume Voices page in generate mode takes a persona prompt of 20 to 2000 characters and makes three variants for you to choose from. Sume's generate mode lives in the web app's Assets area, and the voice you save gets a library id you can reuse in TTS requests.
If you wanted a clone of a real person, that is a different feature; this one invents a voice from words.
What ElevenLabs documents
The ElevenLabs voices doc lists Voice Design with a description of 20 to 1000 characters. It also lists Instant Voice Clone from under two minutes of audio and Professional Voice Clone, which needs a Creator plan or above and can be shared in the Voice Library. The Eleven v4 post separately mentions clones from about 10 seconds, so check which number applies to your plan.
What Sume's Voices page does
In the Voices page of the web app, generate mode asks for a persona prompt and a gender (male, female or nonbinary), and a language chosen from a list of 16: English, Korean, Japanese, Chinese, Spanish, French, German, Portuguese, Italian, Hindi, Dutch, Polish, Russian, Swedish, Turkish and Tagalog. The prompt must be at least 20 and at most 2000 characters, and the page produces three variants to audition. A saved voice carries an id beginning voi_ followed by 32 hex characters, which the TTS contract accepts as voice.id.
| Item | ElevenLabs Voice Design | Sume Voices, generate mode |
|---|---|---|
| Description length | 20 to 1000 characters | 20 to 2000 characters |
| Output | A voice to preview and save | Three variants to choose from |
| Where used | ElevenLabs products and API | Sume TTS via voice.id |
| Clone from a real sample | Separate feature | Separate mode on the same page |
Writing a good description
Say who the speaker is, how they sound and how they should deliver lines. Age range, pace, warmth and setting beat vague adjectives. Avoid naming a real celebrity; you want an original voice that you have the right to use.
- State the use: narration, ad read, tutorial.
- Describe pace and energy in plain words.
- Mention accent only if you have tested it in your target language.
- Pick the language before the persona, since the voice is tied to it.
After you save
Run the same ten-line script through all three variants, and keep the winner's id with a note about the language you tested it in. For delivery tweaks like speed and emotion, use the generation_config fields in the TTS request rather than rewriting the persona.
Sources
Related posts
More in Comparisons
- Does adding references help? A 3 versus 10 reference test plan on Sume
FLUX 3 Image takes ten references. Before you assume more is better, run 3 against 10 on your own products. A scoring sheet and script, no invented results.
- Does Sume have Veo 3.1? No. Which video models it lists instead
Sume's video catalog has no Veo model. See the models it does list, with duration and resolution ranges, and how Google's own docs now steer to Omni Flash.
- Editframe cloud render $0.02 a minute plus $99 vs Sume Timeline
Editframe bills cloud render minutes by megapixel band from $0.02, on a $99 plan. A 1080x1920 Reel is 2.07 MP. Sume renders a minute for $0.10 with no base fee.
- Edits has 250+ fonts: Sume's 29 caption fonts are Hangul-only
Edits ships 250+ fonts and 50+ text animations. Sume Video Captions lists 29 fonts, all Hangul, and no font field for Latin styles. What that means.
Written by Sume