Synthesia voice clone: consent passcode and 1-5 minute upload vs Sume

Synthesia clones a voice from a recording or a 1-5 minute upload and a spoken consent passcode. Sume has no voice-clone route; here is the practical gap.

5 min readSume
All posts

To clone a voice in Synthesia you either record yourself or upload 1 to 5 minutes of audio, then the speaker reads a randomly generated passcode aloud as consent. Sume does not offer a voice-clone route: its avatar talking-video takes script text and speaks it, and the docs I searched contain no clone or consent endpoint.

Synthesia's steps are from its help article How do I clone my voice in Synthesia?, and Sume's from Generate avatar video, both read on 2026-10-02.

What does Synthesia require?

You select Create voice, choose to record or upload, and provide gender, language and microphone or audio details. Recording asks you to read a script in a positive tone with natural pauses. Uploads must be 1 to 5 minutes long in mp3, wav, m4a, weba, aac, flac or ogg.

Consent is recorded by speaking the passcode aloud; a silent recording is rejected and the consent recording must be under 60 seconds. The page says the person whose voice is cloned must be over the applicable legal age and must give consent themselves, not on someone else's behalf. Cloning is included on all subscription plans, though some languages are limited to Enterprise. The finished voice shows under Custom Voices in the voice picker.

What does Sume do about voices?

On POST /v1/avatar-1.0/talking-video each spoken scene carries voice.type: "text" with a script or input_text, and Sume produces the speech and the lip-synced video in the same job. There is no field for uploading a voice sample, no consent step, and no custom-voice list in the docs I read. To see how the voice is chosen for an avatar, read which voice your avatar speaks with.

Side by side

The difference is whether the voice is yours or the platform's.

Voice cloning as documented, read 2026-10-02
QuestionSynthesiaSume
Clone from a recordingYesNo route documented
Sample lengthUpload 1-5 minutesNot applicable
Consent stepSpeak a random passcode, under 60 secondsNot applicable
Formatsmp3, wav, m4a, weba, aac, flac, oggNot applicable
Speech from textYes, in the script boxYes, script or input_text per scene

What should I do if I need my own voice?

If a clone of a specific person is the requirement, Sume is the wrong tool today and Synthesia's flow, with its consent check, fits better. If you only need a consistent, natural voice across many clips, keep one avatar handle and write scripts, which Sume handles per second of video. Do not work around the gap by cloning someone without their consent; Synthesia's own page treats consent as something only the speaker can give, and the same logic applies to any provider. See the HeyGen comparison for another vendor with the same difference.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume