Synthesia voice clone: consent passcode and 1-5 minute upload vs Sume
Synthesia clones a voice from a recording or a 1-5 minute upload and a spoken consent passcode. Sume has no voice-clone route; here is the practical gap.

To clone a voice in Synthesia you either record yourself or upload 1 to 5 minutes of audio, then the speaker reads a randomly generated passcode aloud as consent. Sume does not offer a voice-clone route: its avatar talking-video takes script text and speaks it, and the docs I searched contain no clone or consent endpoint.
Synthesia's steps are from its help article How do I clone my voice in Synthesia?, and Sume's from Generate avatar video, both read on 2026-10-02.
What does Synthesia require?
You select Create voice, choose to record or upload, and provide gender, language and microphone or audio details. Recording asks you to read a script in a positive tone with natural pauses. Uploads must be 1 to 5 minutes long in mp3, wav, m4a, weba, aac, flac or ogg.
Consent is recorded by speaking the passcode aloud; a silent recording is rejected and the consent recording must be under 60 seconds. The page says the person whose voice is cloned must be over the applicable legal age and must give consent themselves, not on someone else's behalf. Cloning is included on all subscription plans, though some languages are limited to Enterprise. The finished voice shows under Custom Voices in the voice picker.
What does Sume do about voices?
On POST /v1/avatar-1.0/talking-video each spoken scene carries voice.type: "text" with a script or input_text, and Sume produces the speech and the lip-synced video in the same job. There is no field for uploading a voice sample, no consent step, and no custom-voice list in the docs I read. To see how the voice is chosen for an avatar, read which voice your avatar speaks with.
Side by side
The difference is whether the voice is yours or the platform's.
| Question | Synthesia | Sume |
|---|---|---|
| Clone from a recording | Yes | No route documented |
| Sample length | Upload 1-5 minutes | Not applicable |
| Consent step | Speak a random passcode, under 60 seconds | Not applicable |
| Formats | mp3, wav, m4a, weba, aac, flac, ogg | Not applicable |
| Speech from text | Yes, in the script box | Yes, script or input_text per scene |
What should I do if I need my own voice?
If a clone of a specific person is the requirement, Sume is the wrong tool today and Synthesia's flow, with its consent check, fits better. If you only need a consistent, natural voice across many clips, keep one avatar handle and write scripts, which Sume handles per second of video. Do not work around the gap by cloning someone without their consent; Synthesia's own page treats consent as something only the speaker can give, and the same logic applies to any provider. See the HeyGen comparison for another vendor with the same difference.
Sources
Related posts
More in Comparisons
- Tavus per-minute video price vs Sume avatar per second
Tavus lists video generation at $1 per minute overage on Starter. Sume avatar video is $11.04 to $33 per minute depending on quality. Dated 2026-10-01.
- TikTok product avatars skip shoes, hats, sunglasses and bracelets
TikTok's Symphony help page lists products its avatars cannot show. What to try for those SKUs, and what Sume's avatar product_image does and does not claim.
- Together AI dynamic rate limits (429, 503) vs Sume rate_limited
Together AI publishes no fixed tiers: limits track live capacity and your recent traffic. How its 429 and 503 map to Sume's rate_limited and queue_full.
- Together AI video API alternative: Sume /v1/videos compared
Together AI's video API is create-then-retrieve with five statuses. How that maps to Sume /v1/videos, and what each side does better.
Written by Sume