What is an AI avatar? What it does and what it can't

An AI avatar is a digital presenter, a face and a voice, that speaks the script you write in a video. What it does, how it's made, and its limits.

5 min readSume
All posts

An AI avatar is a digital presenter: a face and a voice, animated by AI, that speak the script you write in a video, so nobody has to be filmed. Change the words and you get a new video without a new shoot. The same name is also used for AI profile pictures and game characters; this page is about the presenter that talks on screen.

On Sume, an avatar is a reusable identity you create once, from a text prompt, a short profile, or a photo, and then use by its handle in talking videos of 4 to 60 seconds. It makes video files; it is not a live character that chats back. The Sume details come from the Models overview, Create new avatar and Generate avatar video docs and the Sume API reference, read on 2026-09-28. Anything called current behavior is read from Sume's code.

What does an AI avatar do?

It presents. You supply the words, and the avatar says them on camera: a product explainer, a course segment, a UGC-style ad, or the answer to a question customers keep asking.

On Sume that is a two-step workflow. First you create the avatar. Creation is a job, and when it finishes the avatar becomes a reusable resource in your workspace. Then you send a script, or several scenes as video_inputs, together with the avatar's handle. Each talking video is its own job: submit the request, poll or wait for completion, then read the result URL.

What does an AI avatar look like?

Like a person on a video call: a face and shoulders facing the camera, speaking, with a voice to match. How lifelike it looks depends on the tool that made it.

A Sume avatar is a generated person, even when you start from a photo: in current code, a photo avatar is a new image drawn from your picture, not the picture itself. Once the avatar is ready, its preview_image_url shows that person as a public, Sume-hosted image, so you can check the face before you render anything with it. In current code, creating an avatar also makes its voice, and the avatar reports that voice's status: processing, ready, or failed. How do AI avatars work? explains how the face and the voice are made.

How do you make an AI avatar?

Describe the person, give a short profile, or start from a photo: on Sume those are the three inputs of POST /v1/avatar-1.0/generate, and How to create a reusable AI avatar has the full request. API pricing lists creation at $0.95 per avatar, plus a 5.5% agent fee by default. A handle that is already in use returns 409 avatar_handle_taken, so an avatar you want to redo gets a new handle.

Can an AI avatar talk live, or in any language?

Not a Sume avatar, today. Every talking video is rendered as a job. Even mode: sync waits at most 30 seconds, and that ceiling bounds the HTTP wait, not the job, so the avatar can't hold a live conversation.

In current code the talking video also speaks English only, and each avatar's voice is cloned as English. For another language, make the speech with TTS 1.0 and its language field, then lip sync the avatar to that audio with VEED Fabric 1.0. Which languages can an AI avatar speak? walks through it.

What are the limits?

These hold for every Avatar 1.0 talking video. For a longer piece, the docs say to shorten the script or split it into several jobs.

From Generate avatar video and the Sume API reference, read 2026-09-28. The English-only row is current code.
LimitAvatar 1.0 talking video
LengthEstimated at 4–60 seconds, inclusive
PeopleOne avatar per video; scene backgrounds resolve to one shared scene
Shape1:1, 3:4, 9:16, 4:3, 16:9; default 9:16
Resolution720p is the documented resolution
SpeechEnglish only
ConversationNone: each video is a job you poll, not a live stream

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume