How do AI avatars work? Face, voice, and video, step by step
An AI avatar video is assembled from parts: a face, a voice to match it, and motion that makes the face say your script. How each step works on Sume.
AI avatars work by putting a presenter together from parts: a face, a voice that goes with it, and motion that makes the face say your script. Depending on the tool, the face and the voice start from a photo, a recording or a text description, and the result is a video rendered from your words instead of a new shoot.
On Sume today, creating an avatar generates an identity image from your prompt, profile or photo, plus a voice made to suit that face. Each talking video then estimates how long your script takes to say, renders a first frame, turns the speech into short clips, and composes them into one video; captions and music, when you add them, come last. The steps below come from the Generate avatar video and Avatar video previews docs and the Sume API reference, read on 2026-09-28. Anything called current behavior is read from Sume's code.
How are AI avatars made?
Creating an avatar is one job with two results, a face and a voice:
- The face. In current code this is an image-generation step that outputs one 9:16 portrait. A photo avatar is redrawn from your photo with the prompt "Photo of this person"; prompt and profile avatars are drawn from your text or from the
ethnicity,sexandageyou send. Once the avatar is ready, itspreview_image_urlshows the result. - The voice. Also in current code, Sume generates a speaking sample in a voice meant to match the person's look and age, then clones that voice. The avatar reports the voice's own status:
processing,ready, orfailed. - The link between them. In current code, a talking video requested before the voice is ready is refused with
409 avatar_not_ready("Avatar voice is not ready for video generation."). Once the voice isready, TTS 1.0 can also speak with it when you pass the avatar's handle.
How does an avatar talk?
A talking video starts from a ready avatar and a script, or several scenes in video_inputs, and runs in stages:
- Estimate. In current code Sume counts 2.8 words per second to estimate how long the script takes to say, and the plan must land at 4–60 seconds inclusive. Script length for an AI avatar video does the math.
- First frame. A still of the avatar in its scene opens the render; when several scenes share one scene, later scene stills are pose-anchored continuations of that first frame. An avatar video preview runs only this stage, so you can approve the composition before paying for a full render.
- Clips. The speech is rendered as short clips, 4–12 seconds each in current code, spoken in English only.
- Compose. The clips become one composed video: one avatar and one shared scene per video.
Where do captions and music come in?
After the talking head exists, and when you ask for them. Ask on an avatar video preview, which stores caption and package intent and applies it when generate-video runs: in current code, POST /v1/avatar-1.0/talking-video itself refuses captions and package with 400 invalid_request, even though the reference lists them.
When both are set, the server order is: the clean talking head, captions burned onto that clean video, then music, generated from a prompt or taken from a track you link, mixed under it. If captions or music fail, you keep the furthest successful video: captioned when that worked, otherwise clean.
What does each step mean for my video?
Here is what each step means in practice:
| Stage | What Sume makes | What it means for you |
|---|---|---|
| Avatar image | One 9:16 portrait from your prompt, profile or photo | A photo avatar is a new image; check preview_image_url |
| Voice | A voice made to suit the face, then cloned | Videos are refused with 409 avatar_not_ready until voice.status is ready |
| Estimate | 2.8 words per second | Each video must estimate at 4–60 seconds |
| First frame | Stills of the avatar in its scene | A preview shows them before the full render |
| Clips | 4–12 seconds each, composed into one video | One avatar and one shared scene per video |
| Post-production | Captions, then music | A failed stage keeps the furthest successful video |
Sources
Related posts
More in Sume Avatar 1.0
- How long can HeyGen videos be? Limits by plan and API
HeyGen videos can run 1 minute on Free, 30 on Creator and Pro, 60 on Business, with no max on Enterprise; an API avatar request caps at 30 minutes.
- AI avatar prompts: how to describe a talking presenter
An AI avatar prompt for video describes one person: age, look, hair, clothing, and expression. What to write, what Sume adds, and what it can't set.
- Kling API avatar: one image plus audio, 2 to 300 seconds
Kling's API has an Avatar endpoint: one reference image plus a 2–300 second audio track becomes a talking video, billed per second. Inputs and Sume options.
- Difference between lip sync and dubbing: which do you need?
Dubbing replaces the speech in a video; lip sync matches a mouth to the audio. Lip-sync dubbing does both. What each means, and which one you need.
Written by Sume