Voice clone 10 seconds AI: what a short sample gets you on Sume
ElevenLabs says Instant Voice Clones can capture a voice from 10 seconds of audio. On Sume you clone in the app, from an upload or a 30-second recording.

A 10-second voice clone means a very short recording is enough to make a reusable voice. ElevenLabs' Eleven v4 post says Instant Voice Clones can now capture voices with high fidelity from just 10 seconds of audio. On Sume you clone in the app, not through the API: Assets, Voices takes an uploaded file or a recording of up to 30 seconds, and the resulting voice id goes into your text-to-speech calls.
The ElevenLabs claim is from its launch post, read 2026-09-29, and is its own statement. The Sume facts are from the current app code and the TTS schema in the Sume API reference.
Where does a short sample go on Sume?
The Voices page offers to clone a voice or invent one from a prompt, and its create dialog says: “Upload or record audio to clone, or describe a person and we generate the voice.”
| Input | Current behavior |
|---|---|
| File picker | Accepts audio files: wav, mp3, m4a, webm, ogg and mp4 |
| In-app recording | Stops at 30 seconds |
| Sample length rule | No minimum length appears in the app copy or docs read |
| After cloning | Send the voice's id as voice.id on POST /v1/tts-1.0/generate |
Should I record 10 seconds or 30?
The app copy and docs read for this post state no minimum, so nothing there says 10 seconds is enough or not enough. ElevenLabs' 10-second figure is about its own Instant Voice Clones. Record a clean sample of a few sentences in a quiet room, use the recorder's full 30 seconds if you like, and listen to a test line before you use the voice.
Can I clone a voice through the API?
No route creates a clone. The TTS 1.0 selector is the voice id or an avatar reference, and TTS 1.0 has no engine picker: model and model_id are rejected with a 400. The one-time clone step is manual; everything after it can be scripted. Voice cloning API covers that split.
What about permission to clone a voice?
Clone only your own voice, or one whose owner has agreed. The Sume docs read for this post do not set out a consent procedure, so for the rules that apply to your account, read Sume's terms, and for ElevenLabs' rules, its own pages.
Can I get a voice without any sample?
Yes, in the app. Its create dialog also lets you describe a person and generates the voice. That path needs no recording, and the resulting voice is used the same way, by its id.
ElevenLabs' post also says Professional Voice Clones are supported for Eleven v4, which is its longer-sample tier. Sume's docs list no equivalent tier, so check the app for what it offers today.
Sources
Related posts
More in Use cases
- EU AI Act Article 50 AI video labeling requirements for creators
Article 50 applies from 2 August 2026: deployers must tell people about deepfakes and providers must add machine-readable marks. What that leaves to a creator.
- How to extend an AI video past 30 seconds
Sume has no extend parameter on video models. Chain clips: pull a last frame, use it as the next first_frame, then join with Timeline. Vendor limits too.
- Facebook Reels ad specs: 9:16, 1440x2560, H.264 and safe zones
Meta's ads guide for Facebook Reels lists 9:16, 1440x2560, MP4/MOV, 4 GB, H.264 and safe zones of 14% top and 35% bottom. What Sume can and cannot hit.
- Facebook Reels API video length: 3 to 90 seconds, 24 to 60 fps
Meta's Reels publishing docs list 3–90 seconds, 9x16, 1080x1920 recommended, 24–60 fps, closed GOP and 48 kHz AAC. How Sume durations line up.
Written by Sume