Avatar V needs 15 seconds of you; a Sume avatar needs one still
HeyGen Avatar V clones you from a 15-second webcam clip. Sume builds a reusable avatar from one photo for $0.95. What each needs and what you can render.
HeyGen Avatar V needs a 15-second webcam recording of you and gives back a digital twin. Sume Avatar 1.0 needs one reference photo, or a text prompt, and creates a reusable avatar for a one-time $0.95; it does not learn gestures from a recording, so pick it when a still and a script are all you have.
What HeyGen says it needs
The Avatar V page says '15 seconds to create your avatar' from a short webcam clip. It lists 175+ languages and dialects with phoneme-level lip sync, a G2 ranking as the most realistic avatars, and consistent character across scenes. It says there is no cap on video length or quality. No price appears on the page; it points to a separate pricing page.
The page does not describe any numbers for training time, and we do not guess them here.
What Sume needs
Sume has three ways to make an avatar: a text prompt, structured traits (ethnicity, sex, age), or a reference image. The image must be a fetchable public HTTPS URL. If you only have a recording, export one clear frame from it and use that as the photo.
Each request is a job. When it completes you get an avatar handle, which you pass to POST /v1/avatar-1.0/talking-video with a script.
| Question | HeyGen Avatar V | Sume Avatar 1.0 |
|---|---|---|
| Input to create | 15-second webcam recording | Photo, prompt or traits |
| Learns movement from video? | Yes, per the page | No, it starts from a still |
| Creation price | Not on the page | $0.95 once per avatar |
| Video length per job | No cap stated | 4 to 60 seconds |
| Reuse | Across scenes | Same handle in every job |
Cost of the second video onward
After creation, you pay for seconds of video. At plus quality with no product image the rate is $0.245 per second, so a 45-second clip is 45 x $0.245 = $11.03 (rounded). Ten such clips plus the avatar is 10 x $11.025 + $0.95 = $111.20. The same avatar handle is used in all ten.
When each fits
- You want your own movement and voice captured: a recording-based twin.
- You want a presenter who is not a specific real person, or only have a photo: a Sume avatar.
- You need to run it from a script or spreadsheet: Sume exposes it as an API route with Idempotency-Key.
- You need videos longer than 60 seconds: split the script across jobs.
What to do
Create the avatar with a still, then generate a 10-second test clip at standard quality first: 10 x $0.184 = $1.84. If the face and voice read right, move up to plus.
Sources
Related posts
More in Comparisons
- HeyGen Avatar Shots with Seedance 2.0 vs Sume avatar plus Seedance 2.5
HeyGen Avatar Shots place your digital twin in Seedance 2.0 scenes. On Sume, a talking avatar clip and a Seedance 2.5 shot are separate jobs you combine.
- How many references can a holiday ad take? Five video models compared
Seedance 2.5 lists 30 images, Wan 3.0 up to 20 assets, Omni Flash 1.1 on Sume 10 images and 3 short clips, Luma Ray3.2 16 keyframes. What it means for ad kits.
- Hume plan ladder: effective cost per 1,000 characters vs Sume
Hume's plans include 30K to 10M characters. Used in full, the effective rate runs $0.05 to $0.10 per 1K. Sume TTS is a flat $0.0475. Break-even by plan.
- Ideogram 4.5 vs Seedream 4.5 for image edits: limits and price on Sume
Both edit via POST /v1/images. Ideogram 4.5 takes 5 references, priced by quality; Seedream 4.5 takes 10 at one rate. Catalog facts, read 2026-10-05.
Written by Sume