HeyGen Avatar V's 15-second recording vs a Sume photo avatar
HeyGen's Avatar V starts from a 15-second recording. Sume's avatar API starts from a prompt, traits or one public HTTPS image, with no recording step.
HeyGen's Avatar V starts from a short recording of you; Sume's avatar API takes no recording. You create an avatar from a text prompt, structured traits or one reference image at a public HTTPS URL, then reuse its avatar_handle. The two do not produce the same thing, and this post does not rank them.
What does Avatar V need?
HeyGen's announcement says "Record a 15-second clip" and that "one short recording from you generates studio-quality video that maintains your face, your voice, and your presence across angles, looks, and runtime". Those are HeyGen's own claims.
What does a Sume avatar need?
Per the avatar docs, a request has a top-level avatar_handle plus an input of one of three types. The handle may start with @; Sume stores it without. The job is polled, then the handle is used on avatar videos.
| Input | HeyGen Avatar V | Sume |
|---|---|---|
| A recording | 15-second clip | Not an input in the docs |
| A photo | Not stated in the post | input.type "photo" with image_url |
| Traits | Not stated in the post | input.type "props" (ethnicity, sex, age) |
| A text description | Not stated in the post | Prompt input |
What does the photo route look like?
image_url must be fetchable, public and HTTPS; localhost, private-network, non-HTTPS URLs and non-image responses are rejected before generation.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: avatar-photo-001" \
-d '{
"avatar_handle": "reference_presenter",
"input": {
"type": "photo",
"image_url": "https://example.com/reference.png"
}
}'Does Sume capture my voice?
The avatar docs read for this post describe no voice capture in avatar creation; the avatar video's voice is set by the script or scene voice fields. Use a real person's photo only with their agreement, since the docs show no consent step.
Sources
Related posts
More in Sume Avatar 1.0
- HeyGen Edit Look: retouching an AI avatar, and the Sume route
HeyGen's Edit Look retouches an existing avatar in place. Sume's docs show no such edit, so the route is a new avatar from a retouched photo and a new handle.
- Shortest video an AI dub or face swap accepts: 5 s vs 4 s
Synthesia's dubbing page says a video must be at least 5 seconds. Sume's Beta face swap plans for about 4-15 seconds and avatar videos take 4-60. Side by side.
- Does a face-swapped video need a YouTube AI disclosure?
YouTube asks for disclosure when content makes a real person appear to say or do something they didn't. What that means for a Sume Beta face-swap output.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume