ElevenLabs Professional Voice Clone: own voice only, 24 h retry
ElevenLabs Professional Voice Cloning only accepts your own voice, needs 30 minutes of audio and a verification step. What Sume's Voices library asks instead.

ElevenLabs allows a Professional Voice Clone only of your own voice: its docs say that even with someone's consent you cannot clone another person's voice. Before training you pass a verification step, and if it fails you wait 24 hours or contact support, per the Professional Voice Cloning page (read 2026-10-02).
Sume has no verification step and no professional-clone tier; you clone from a short clip in the app.
The ElevenLabs rules
The quickstart lists eight steps, from creating the voice and uploading audio to speaker separation, a CAPTCHA verification where the voice owner reads text aloud, and training. It says verification is required to confirm you have permission to use the voice.
| Item | Detail |
|---|---|
| Whose voice | Your own only, even with consent for another person |
| Minimum audio | 30 minutes of high-quality audio |
| Recommended audio | Closer to 2 to 3 hours |
| Verification retry | Wait 24 hours or contact support |
| Plan slots | None on Free and Starter; 1 on Creator, Pro and legacy Scale |
| Training time | Usually 3 to 6 hours after verification |
| Languages | All languages in the Eleven v4 family, 90+ |
What Sume offers
Sume's Voices library in the app has a create dialog that takes a name, a gender, a language and an uploaded or microphone-recorded clip. That dialog has no consent-recording step, and the public API has no route that creates a clone.
There is no slot count in the dialog and no multi-hour training queue to wait for. The trade is control: you cannot ask Sume for a 3-hour fine-tune of a voice.
How this changes a project plan
The vendor rules change who can be the narrator.
- Own-voice only is an ElevenLabs rule for PVC. A narrator who is not you must use another route there; on Sume the responsibility to hold the release is yours.
- For a branded voice that must sound identical across a year of content, decide on one stored clip and keep it.
- Narration for video: use the voice id in text to speech, then join the lines with Timeline audio.
Related reading
For the 10-second instant clone side of the same vendor, see Voice clone 10 seconds AI. For the clone route on Sume, see Voice cloning API.
Sources
Related posts
More in Comparisons
- ElevenLabs use_pvc_as_ivc: what it does and Sume's voice selector
ElevenLabs' use_pvc_as_ivc flag swaps a professional voice clone for its instant version. Sume's TTS has no such switch: you pick an avatar or a voice id.
- ElevenLabs stability 0.5 and similarity 0.75 vs Sume generation_config
ElevenLabs defaults stability to 0.5 and similarity to 0.75. Sume has no such sliders: generation_config takes only volume, speed and an emotion string.
- fal Agent (early access) vs Sume Agent Completions and hosted MCP
fal's August 2026 agent is early access, with API, CLI and MCP. Sume's agent has Agent Completions over the API and a hosted MCP. What each documents today.
- fal cancel returns 202 or 400: how Sume's cancel 409 differs
fal's queue cancel answers 202 CANCELLATION_REQUESTED, 400 ALREADY_COMPLETED or 404. Sume's cancel works only before generation starts, else 409.
Written by Sume