HeyGen Avatar V's 15-second recording vs a Sume photo avatar

HeyGen's Avatar V starts from a 15-second recording. Sume's avatar API starts from a prompt, traits or one public HTTPS image, with no recording step.

4 min readSume
All posts

HeyGen's Avatar V starts from a short recording of you; Sume's avatar API takes no recording. You create an avatar from a text prompt, structured traits or one reference image at a public HTTPS URL, then reuse its avatar_handle. The two do not produce the same thing, and this post does not rank them.

What does Avatar V need?

HeyGen's announcement says "Record a 15-second clip" and that "one short recording from you generates studio-quality video that maintains your face, your voice, and your presence across angles, looks, and runtime". Those are HeyGen's own claims.

What does a Sume avatar need?

Per the avatar docs, a request has a top-level avatar_handle plus an input of one of three types. The handle may start with @; Sume stores it without. The job is polled, then the handle is used on avatar videos.

Avatar creation inputs, read 2026-09-30 (Sume avatar docs, HeyGen).
InputHeyGen Avatar VSume
A recording15-second clipNot an input in the docs
A photoNot stated in the postinput.type "photo" with image_url
TraitsNot stated in the postinput.type "props" (ethnicity, sex, age)
A text descriptionNot stated in the postPrompt input

What does the photo route look like?

image_url must be fetchable, public and HTTPS; localhost, private-network, non-HTTPS URLs and non-image responses are rejected before generation.

curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-photo-001" \
  -d '{
    "avatar_handle": "reference_presenter",
    "input": {
      "type": "photo",
      "image_url": "https://example.com/reference.png"
    }
  }'

Does Sume capture my voice?

The avatar docs read for this post describe no voice capture in avatar creation; the avatar video's voice is set by the script or scene voice fields. Use a real person's photo only with their agreement, since the docs show no consent step.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume