How to create an avatar from a video: two ways

To create an avatar from a video, train a digital twin on the footage, or take one clear frame and make a photo avatar from it. What each route keeps.

5 min readSume
All posts

To create an avatar from a video, you either give the footage to a tool that trains a digital twin from it, or you take one clear frame of the person and make a photo avatar from that still. Training uses the whole recording (HeyGen also clones the voice from it); the frame route works with any avatar API that takes an image, including Sume's.

This page is about talking avatars of a real person, not cartoon filters. Vendor facts come from HeyGen's and Synthesia's own docs, read on 2026-09-28; Sume facts come from Create new avatar and Sume's code.

How do I train an avatar from my footage?

Upload the recording to a product that builds avatars from video. Each sets its own rules for the footage:

From HeyGen's Video to Avatar and API pricing pages and Synthesia's create an avatar page, read 2026-09-28.
RuleHeyGen digital twinSynthesia avatar from video (legacy)
FootageTargets: 15-600 seconds; one person facing the camera; clear speech throughout1-5 minutes in a single continuous take; .webm, .mp4 or .mov up to 2GB
VoiceOne voice cloned from the same footage automaticallyVoice cloning built into the flow
ConsentThe subject records a short statement through a link valid for 24 hoursA consent video recorded live by the same person
Access and timingHeyGen's API pricing article: creating a custom digital twin through the API is for Enterprise API usersReady in 24 hours

How do I make an avatar from one frame of the video?

Sume's avatar API takes no video: POST /v1/avatar-1.0/generate accepts a prompt, a profile, or a reference image. From a recording, the image is your starting point:

  • Pause the video on a frame where the face is sharp, front-facing and evenly lit, with only that person in shot, and export it as an image.
  • Put the image at a public HTTPS URL. Localhost, private-network, non-HTTPS and non-image URLs are rejected before generation.
  • Send it as a photo input with a handle you will reuse, then poll the job until it is completed.
curl -X POST https://api.sume.com/v1/avatar-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: avatar-from-frame-001" \
  -d '{
    "avatar_handle": "reference_presenter",
    "input": {
      "type": "photo",
      "image_url": "https://example.com/frame.png"
    }
  }'

What carries over from the video, and what doesn't?

The frame route keeps only the look in that one frame. In Sume's current code:

  • The likeness is generated, not copied: the photo is sent to an image model with the prompt "Photo of this person".
  • The voice is not yours. Current code renders a sample from the avatar's image and clones that voice in English.
  • Nothing else from the recording is used, since the create call has no video input.
  • To speak in your own voice, clone it in the Sume app (Assets → Voices), where you upload or record audio, and lip sync the frame to it; clone yourself with AI covers that route.

Can I send the video file to Sume directly?

No. Sume has no public upload route for local files, and its API reference says to prefer public HTTPS media URLs in generation requests. The avatar create call takes an image, not a video, so host the frame yourself, as above.

Creating the avatar costs $0.95 per avatar once, plus a 5.5% agent fee by default. The handle then works on every talking video; create a reusable AI avatar covers the rest of the create call, and talking avatar video API covers the next step.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume