Syren can't use your own avatar or cloned voice; Sume has a handle

Synthesia says own avatars and cloned voices need its regular editor, not Syren. Sume builds an avatar from a photo; voice cloning is app-only.

4 min readSume
All posts

On Synthesia's Syren page, the limit is stated plainly: to use your own Personal Avatar or cloned voice, create your video in Synthesia Studio instead (read 2026-10-11). On Sume you can build a reusable avatar from your own reference photo and generate video with its handle through the API, but cloning your own voice is an app feature, not an API one.

The Syren fact is from Introducing Syren: On-Brand AI Video, read 2026-10-11. The Sume facts are from Create new avatar and Generate avatar video.

What Syren does and does not take

Syren includes Synthesia's avatars and voices and lets you swap them. A personal avatar or cloned voice you made yourself is not usable in a Syren video, according to the page. If your brand video must show you or a named spokesperson, that pushes the work back to Synthesia's regular editor.

What Sume does for your own face

An avatar is created from one of three inputs: a text prompt, structured traits, or a reference image sent as input.type: "photo" with a public HTTPS image_url. Creation is a flat $0.95 per avatar. The avatar_handle you choose is stored without a leading @, and every later avatar video request names the handle. Video from that handle is script-driven, 4 to 60 seconds, English only.

Get the person's written consent before uploading a photo of a real person, and keep a record keyed by handle. The corpus already has a release checklist for it.

A few input rules save retries. The image must be a fetchable public HTTPS URL: Sume rejects localhost, private-network URLs, non-HTTPS URLs and non-image responses before it submits generation. Handles are validated too, so check yours before calling. Use an Idempotency-Key on creation so a retry after a timeout does not charge the $0.95 twice.

Who can do what

The table shows where each tool stands for a spokesperson who is a real person.

Own-likeness support as documented (Syren page read 2026-10-11; Sume docs)
NeedSyrenSume
Your face as the presenterNot in Syren; use Synthesia's regular editorAvatar from a reference photo, reuse the handle
Your cloned voiceNot in SyrenCloning is in the Sume app only; not in the API
Stock presenterYes, avatars and voices includedAvatar from a prompt or traits
Reuse across videosWithin Synthesia's editorOne handle for every request
Cost to createNot stated on the page$0.95 per avatar

The voice gap, and what to do about it

If the voice must be yours, there is no API path to a cloned voice on Sume today. The avatar speaks with the voice that comes with the avatar when it reads your script, and the docs do not describe a way to hand it your own recording in the avatar-video route. Sume's separate talking-still route accepts a Sume-hosted audio file, but that is a different product from Avatar 1.0 and it limits the audio to Sume-hosted URLs.

So for a founder video in the founder's own voice, the honest options are to use Synthesia's regular editor, or to accept a stock voice from the avatar. For a brand character with a consistent but invented voice, the avatar handle is the simpler route.

One more practical point: a likeness is a legal and platform question as well as a technical one. A real person's face needs their permission regardless of the tool, and a realistic synthetic person may need a disclosure line where you publish it. Settle consent and disclosure first, then pick the route that your approval allows.

If you are choosing between the two for a founder-led channel, ask which matters more: that the video is built from a chat prompt with brand motion, or that the presenter is demonstrably you. Syren's page puts the first inside Syren and the second outside it. Sume puts the second inside the API for the face and outside it for the voice.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume