Record a voice sample in the browser: Sume Voices 30 seconds
Sume Voices clone mode takes an uploaded clip or a browser recording of up to 30 seconds. ElevenLabs states clone samples from 10 seconds to two minutes.

In the Sume web app, the Voices page has a clone mode where you either upload an audio clip or record one in the browser, and the in-browser recorder stops at 30 seconds. This is a web app feature; the public API documents no cloning endpoint, so a pipeline that must clone programmatically cannot do it on Sume today.
ElevenLabs quotes different sample lengths for its own clones, so compare the numbers with the right label.
What ElevenLabs says about sample length
The ElevenLabs voices doc says Instant Voice Clone needs less than two minutes of audio, and that Professional Voice Clone requires a Creator plan or above and can be shared in the Voice Library. The Eleven v4 post says Instant Voice Clone works from 10 audio seconds. The two pages disagree, so budget a longer, clean sample and treat 10 seconds as the floor, not the target.
| Source | Statement |
|---|---|
| ElevenLabs voices doc | Instant clone: less than two minutes of audio |
| ElevenLabs Eleven v4 post | Instant Voice Clone from 10 audio seconds |
| Sume web app, Voices page | Browser recorder stops at 30 seconds; a file can be uploaded instead |
Recording a good sample
Quality of the take matters more than length. Record in a quiet room, keep the microphone at a steady distance, speak in the style you want the voice to deliver, and avoid music or other voices in the background.
- Use your own voice or one you have written permission to use.
- Keep the language of the sample the same as the language you will synthesize.
- If 30 seconds is not enough variety, upload a longer clean clip instead.
- Re-record rather than editing out noise.
After you have a voice
A saved voice gets a library id beginning voi_ followed by 32 hex characters, which the TTS contract accepts as voice.id. Set the language on every request; the mismatch warning appears when a voice and language disagree, and you confirm with confirm_language_mismatch only after a person agrees.
Because cloning is a web-app flow, record the voice id in your project notes so a script or agent can use it later.
Limits to keep in mind
The permission to clone a voice is your responsibility. Cloning a voice you do not have the right to use creates legal and platform risks, regardless of how short the sample is. See the earlier cloning post for the API-side picture.
Sources
Related posts
More in Comparisons
- Replicate MCP discovery via server.json vs Sume's MCP URL
Replicate publishes /.well-known/mcp/server.json for the official MCP Registry. Sume documents one hosted MCP URL and OAuth metadata. How each client connects.
- Replicate allow_fallback_model: Nano Banana Pro vs Sume
Replicate can fall back from Nano Banana Pro to Seedream 5.0 lite and bills the fallback. Sume's allow_fallbacks is accepted but has no effect. What to do.
- Replicate predictions source=web filter vs Sume jobs list
Replicate lets you list only web-created predictions, limited to 14 days. Sume's GET /v1/jobs lists only jobs your key's member created. How to scope a list.
- Rev AI human transcription $1.99 a minute vs Sume STT $0.01 a minute
Rev AI lists human transcription at $1.99 a minute, 199 times Sume's machine STT at $0.01. What the price gap pays for and when machine output is enough.
Written by Sume