Synthesia personal avatar vs studio avatar: what differs
A Synthesia personal avatar is self-serve, from one photo, ready in minutes; a studio avatar is a green-screen shoot sold as a paid add-on.
A Synthesia personal avatar is a self-serve digital twin you make in the Synthesia app: the recommended version is built from a single photo on Synthesia's Express-2 model and is typically ready within minutes of the consent step. A studio avatar is a produced shoot: three 2–3 minute takes on green screen to Synthesia's filming spec, a consent recording, and its Express-1 model, sold as a paid add-on.
The Synthesia facts come from its own help docs, Create an Avatar, Studio Avatars and Avatars, and its pricing page, read on 2026-09-28. Synthesia lists three custom types: personal avatar from photo, personal avatar from video (legacy), and studio avatar.
What is the difference between a personal and a studio avatar?
Mostly the input and the effort. A personal avatar comes from what you already have (a photo, or older webcam footage). A studio avatar needs a shoot to Synthesia's spec, with a high-quality camera and an external microphone.
| Feature | Personal, from photo | Personal, from video (legacy) | Studio avatar |
|---|---|---|---|
| Built from | One photo, uploaded or taken with your webcam | A webcam recording of a script, or 1–5 minutes of uploaded footage in one continuous take | Three 2–3 minute takes on green screen, plus a consent take |
| Avatar model | Express-2 | Looping auto-alignment | Express-1 |
| Outfits, spaces, poses | Yes (customizable) | Not listed as customizable: reuses your appearance and background | No: Express-1 avatars are not customizable |
| Ready in | Typically minutes after the consent step | Next business day (the creation steps say 24 hours) | Up to 10 days to process |
| Who can make one | Paid plans only | Not stated on the pages read | Paid add-on, $1,000/year, annual plans only; performer 18+ |
How do I create a personal avatar in Synthesia?
On the Avatars page, select + Create avatar, then Create Your Personal Avatar > Create my Avatar > Avatar from Photo. Then:
- Upload a photo of you and only you, or use your webcam. Synthesia suggests landscape orientation, smiling, framed waist up, with your teeth visible for lip sync.
- Optionally record or upload a voice sample to clone your voice; skip it and the avatar uses a Synthesia voice.
- Record the consent video when prompted, then submit.
What does a Synthesia studio avatar shoot need?
Synthesia's filming spec is strict, and it warns a reshoot might be needed if it isn't met:
- Three videos reading the performance script, plus one reading the consent script in the performer's native language. The performer must be at least 18.
- MP4, H.264, under 2GB per video, at 29.97 or 30 fps. Minimum 1920 × 1080; 4K UHD preferred.
- Green screen (blue if the performer wears green), three-point lighting, a lavalier or boom mic, and no webcams or streaming software.
- No jump cuts or edits that remove parts of the performance. The audio is also used to create a voice clone.
Can I use my personal avatar through the Synthesia API?
Synthesia says so: the Avatars page lets you copy an avatar's ID, which it calls handy for referencing that avatar from the Synthesia API, and your personal avatars are listed there under My Avatars. Its pricing page lists API access and 5 personal avatars on the Creator plan, and unlimited personal avatars on Enterprise. Credits and per-minute costs are in Synthesia API pricing and credits.
Can I create a custom avatar from a photo by API instead?
Yes, on Sume. POST /v1/avatar-1.0/generate with an input of type photo and a public HTTPS image_url creates a reusable avatar as a job; it costs $0.95 per avatar. You then send its avatar_handle and a script to POST /v1/avatar-1.0/talking-video for 4–60 second clips. Avatar speech is English only in current code. Details are on Create new avatar and Generate avatar video; the video-footage route is in How to create an avatar from a video.
Sources
Related posts
More in Sume Avatar 1.0
- Talking head vs B-roll: what's the difference?
A talking head is the shot of someone speaking to camera; B-roll is the footage you cut to while their voice keeps playing. How the two fit together.
- UGC video prompt for AI avatars: what goes in each field
A UGC video prompt for an AI avatar is three inputs, not one: the creator's words, a scene prompt for the setting and light, and a product image.
- What is a talking head video? Meaning and how to make one
A talking head video shows one person speaking to the camera, framed from about the chest up. What it means, what it's for, and how AI makes one.
- What is AI UGC? Meaning, how it's made, and how it differs
AI UGC is UGC-style video made with AI: a generated presenter speaks your script to camera, the way a creator would on a phone. How it differs.
Written by Sume