Avatar V multi-look vs Sume scene prompts
HeyGen Avatar V adds multi-look generation. On Sume an avatar keeps one identity, and each video's look comes from scene prompts, product images and quality.
On Sume, a reusable avatar is one identity, created from a prompt, structured traits or a reference image. The look of a given video is steered per video with a scene prompt or photo, an optional product image and a quality tier. The docs describe no outfit-swapping control, so a different look means different scene direction or a separate avatar.
What HeyGen announced
HeyGen's announcement, dated Sep 17, 2026, describes Avatar V as producing studio-quality video from a 15-second phone recording, with multi-angle stability and multi-look generation. It lists 177+ languages and dialects. The page does not state pricing, so no price is compared here.
How Sume creates an avatar
POST /v1/avatar-1.0/generate takes a top-level avatar_handle and an input union with three types: prompt (describe the avatar), props (structured traits such as ethnicity, sex and age) and photo (a public HTTPS image_url). The handle may include a leading @; Sume stores it without it. The input is an image, not a phone video.
How a video's look is steered
POST /v1/avatar-1.0/talking-video references the ready avatar by avatar_handle and takes either a script or ordered video_inputs. Planned duration must land between 4 and 60 seconds.
scene: { "type": "prompt", "prompt": "..." }sets scene direction;scene: { "type": "photo", "image_url": "..." }gives a photo reference.product_imageis optional and adds a product to the video.qualityisstandard(fastest),plus(default) ormax(highest, slower).aspect_ratiosupports 1:1, 3:4, 9:16, 4:3 and 16:9, default 9:16.
Side by side
Both rows are limited to what each page states.
| Topic | HeyGen Avatar V | Sume Avatar 1.0 |
|---|---|---|
| Source for the avatar | 15-second phone recording | Prompt, traits or one reference image |
| Different looks | Multi-look generation | Scene prompt or photo per video; separate avatar per identity |
| Longest script | Not stated on the page | 4 to 60 seconds per video |
| Pricing | Not stated on the page | Per-job tiers; check the catalog |
Current execution supports one resolved avatar per final video, with scene backgrounds resolving to one shared scene. Plan a campaign of different looks as several jobs rather than one.
Sources
Related posts
More in Sume Avatar 1.0
- Italy deepfake offense: 1-5 years, and avatar consent records
A secondary source says Italy's Law 132/2025 took effect Oct 10, 2025 and Art. 612-quater carries 1-5 years for deepfakes. Keep consent records for avatars.
- Split a 3-minute script into 60-second avatar jobs
Sume avatar talking videos accept scripts of an estimated 4-60 seconds. Split a 3-minute script into jobs of that size, then join the audio with timeline audio.
- Udemy promo video at 90 s: an avatar intro in two clips
Udemy's guidance puts an ideal promo video near 90 seconds. A Sume avatar clip tops out at 60 seconds, so build the intro as two clips and join them.
- Alt text for an avatar video poster: WCAG 1.1.1 in practice
A poster image from a Sume avatar preview still needs a text alternative under WCAG 1.1.1, and the video needs descriptive identification. What to write.
Written by Sume