How do AI avatars work? Face, voice, and video, step by step

An AI avatar video is assembled from parts: a face, a voice to match it, and motion that makes the face say your script. How each step works on Sume.

5 min readSume
All posts

AI avatars work by putting a presenter together from parts: a face, a voice that goes with it, and motion that makes the face say your script. Depending on the tool, the face and the voice start from a photo, a recording or a text description, and the result is a video rendered from your words instead of a new shoot.

On Sume today, creating an avatar generates an identity image from your prompt, profile or photo, plus a voice made to suit that face. Each talking video then estimates how long your script takes to say, renders a first frame, turns the speech into short clips, and composes them into one video; captions and music, when you add them, come last. The steps below come from the Generate avatar video and Avatar video previews docs and the Sume API reference, read on 2026-09-28. Anything called current behavior is read from Sume's code.

How are AI avatars made?

Creating an avatar is one job with two results, a face and a voice:

  • The face. In current code this is an image-generation step that outputs one 9:16 portrait. A photo avatar is redrawn from your photo with the prompt "Photo of this person"; prompt and profile avatars are drawn from your text or from the ethnicity, sex and age you send. Once the avatar is ready, its preview_image_url shows the result.
  • The voice. Also in current code, Sume generates a speaking sample in a voice meant to match the person's look and age, then clones that voice. The avatar reports the voice's own status: processing, ready, or failed.
  • The link between them. In current code, a talking video requested before the voice is ready is refused with 409 avatar_not_ready ("Avatar voice is not ready for video generation."). Once the voice is ready, TTS 1.0 can also speak with it when you pass the avatar's handle.

How does an avatar talk?

A talking video starts from a ready avatar and a script, or several scenes in video_inputs, and runs in stages:

  • Estimate. In current code Sume counts 2.8 words per second to estimate how long the script takes to say, and the plan must land at 4–60 seconds inclusive. Script length for an AI avatar video does the math.
  • First frame. A still of the avatar in its scene opens the render; when several scenes share one scene, later scene stills are pose-anchored continuations of that first frame. An avatar video preview runs only this stage, so you can approve the composition before paying for a full render.
  • Clips. The speech is rendered as short clips, 4–12 seconds each in current code, spoken in English only.
  • Compose. The clips become one composed video: one avatar and one shared scene per video.

Where do captions and music come in?

After the talking head exists, and when you ask for them. Ask on an avatar video preview, which stores caption and package intent and applies it when generate-video runs: in current code, POST /v1/avatar-1.0/talking-video itself refuses captions and package with 400 invalid_request, even though the reference lists them.

When both are set, the server order is: the clean talking head, captions burned onto that clean video, then music, generated from a prompt or taken from a track you link, mixed under it. If captions or music fail, you keep the furthest successful video: captioned when that worked, otherwise clean.

What does each step mean for my video?

Here is what each step means in practice:

From Generate avatar video, Avatar video previews and the Sume API reference, read 2026-09-28. The portrait, voice, 2.8 words-per-second and clip-length details are current code.
StageWhat Sume makesWhat it means for you
Avatar imageOne 9:16 portrait from your prompt, profile or photoA photo avatar is a new image; check preview_image_url
VoiceA voice made to suit the face, then clonedVideos are refused with 409 avatar_not_ready until voice.status is ready
Estimate2.8 words per secondEach video must estimate at 4–60 seconds
First frameStills of the avatar in its sceneA preview shows them before the full render
Clips4–12 seconds each, composed into one videoOne avatar and one shared scene per video
Post-productionCaptions, then musicA failed stage keeps the furthest successful video

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume