What is an AI digital twin? Twins vs photo and stock avatars
An AI digital twin is a custom avatar of one real person that speaks new scripts in their likeness. HeyGen trains it on footage; Synthesia starts from a photo.

An AI digital twin is a custom AI avatar of one real person that can be made to speak new scripts in their likeness without filming again, with a clone of their own voice when the vendor makes one. HeyGen uses the name for an avatar trained on video footage of the person; Synthesia uses it for its Personal Avatars, which it now recommends building from a single photo.
In engineering, "digital twin" also means a virtual model of a machine or system; this page is about the AI video meaning. Vendor facts come from HeyGen's and Synthesia's own docs, read on 2026-09-28, and Sume facts from Create new avatar and Sume's code.
What kinds of AI avatars are there?
Vendors use different names, but the avatar types differ mainly in what the avatar is made from:
| Type | Made from | Vendor example |
|---|---|---|
| Digital twin | Video of one real person, with the voice cloned from the same recording | HeyGen Digital Twin |
| Personal avatar from video | A recording of the person; Synthesia now labels this route legacy | Synthesia |
| Studio avatar | Three recordings of 2-3 minutes each, filmed on green screen; a $1,000-a-year paid add-on for annual plans | Synthesia Studio (EXPRESS-1) |
| Photo avatar | One still image of a person, animated to the script | HeyGen Photo Avatar; Synthesia personal avatar from photo |
| Stock avatar | A ready-made presenter the vendor provides | Both vendors' avatar libraries |
Is a digital twin the same as a photo avatar?
No. A digital twin is trained on video of the person; a photo avatar has only one frame to work from. Synthesia says it plainly: its photo avatars use a speech-driven animation model rather than motion learned from recorded footage, so the avatar's motion will not exactly match how you move in real life.
Voice follows the same split. HeyGen's digital twin clones one voice from the training footage automatically. Synthesia's photo route makes voice cloning optional; if you skip it, the avatar uses a Synthesia voice.
Is consent required for a digital twin?
Yes, from the person being cloned, and a vendor checks it before the twin can be used. The recording rules, consent steps and turnaround for each vendor are in how to create an avatar from a video. Two limits sit outside that flow:
- Age: Synthesia's Studio Avatars page says you must be at least 18 to create one.
- Likeness: a photo avatar skips the consent step in HeyGen's API, but HeyGen's consent page says that is not permission to use someone's likeness.
Does Sume make digital twins?
No. Sume's avatars are generated likenesses, not twins of a real person. POST /v1/avatar-1.0/generate takes one of three inputs: a prompt, a profile of traits, or a reference image at a public HTTPS URL. There is no video input, and even from a reference image the face and the voice are generated rather than copied.
- For ready-made presenters,
POST /v1/avatar-catalog/searchsearches Sume's stock catalog; see stock AI avatars. - To pair your own face and voice anyway, clone yourself with AI covers the lip sync route.
Sources
Related posts
More in Sume Avatar 1.0
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
Written by Sume