Avatar V multi-look vs Sume scene prompts

HeyGen Avatar V adds multi-look generation. On Sume an avatar keeps one identity, and each video's look comes from scene prompts, product images and quality.

4 min readSume
All posts

On Sume, a reusable avatar is one identity, created from a prompt, structured traits or a reference image. The look of a given video is steered per video with a scene prompt or photo, an optional product image and a quality tier. The docs describe no outfit-swapping control, so a different look means different scene direction or a separate avatar.

What HeyGen announced

HeyGen's announcement, dated Sep 17, 2026, describes Avatar V as producing studio-quality video from a 15-second phone recording, with multi-angle stability and multi-look generation. It lists 177+ languages and dialects. The page does not state pricing, so no price is compared here.

How Sume creates an avatar

POST /v1/avatar-1.0/generate takes a top-level avatar_handle and an input union with three types: prompt (describe the avatar), props (structured traits such as ethnicity, sex and age) and photo (a public HTTPS image_url). The handle may include a leading @; Sume stores it without it. The input is an image, not a phone video.

How a video's look is steered

POST /v1/avatar-1.0/talking-video references the ready avatar by avatar_handle and takes either a script or ordered video_inputs. Planned duration must land between 4 and 60 seconds.

  • scene: { "type": "prompt", "prompt": "..." } sets scene direction; scene: { "type": "photo", "image_url": "..." } gives a photo reference.
  • product_image is optional and adds a product to the video.
  • quality is standard (fastest), plus (default) or max (highest, slower).
  • aspect_ratio supports 1:1, 3:4, 9:16, 4:3 and 16:9, default 9:16.

Side by side

Both rows are limited to what each page states.

Avatar V and Sume Avatar 1.0 (read 2026-10-03)
TopicHeyGen Avatar VSume Avatar 1.0
Source for the avatar15-second phone recordingPrompt, traits or one reference image
Different looksMulti-look generationScene prompt or photo per video; separate avatar per identity
Longest scriptNot stated on the page4 to 60 seconds per video
PricingNot stated on the pagePer-job tiers; check the catalog

Current execution supports one resolved avatar per final video, with scene backgrounds resolving to one shared scene. Plan a campaign of different looks as several jobs rather than one.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume