What is a talking head video? Meaning and how to make one
A talking head video shows one person speaking to the camera, framed from about the chest up. What it means, what it's for, and how AI makes one.

A talking head video shows one person speaking straight to the camera, framed from about the chest up, with their words carrying the video. Explainers, announcements, lessons, vlogs and UGC-style ads can all be talking head videos. You can film one, or make one without a camera: an AI avatar that speaks your script, or a still photo lip-synced to recorded or generated speech.
On Sume those are two routes: an Avatar 1.0 talking video, 4 to 60 seconds long, made from a script and a ready avatar; or VEED Fabric 1.0, which animates one still image or avatar to Sume-hosted audio of up to 300 seconds. The Sume facts come from the Generate avatar video and Models overview docs and the Sume API reference, read on 2026-09-28; the definitions are general. Anything called current behavior is read from Sume's code.
What does talking head mean in video?
It names the shot: a head and shoulders, talking. In an interview, the talking head is the person answering; in a presenter video, it is the host. In editing, the talking head is the main footage of the speaker, also called A-roll, and the footage you cut away to while the voice keeps going is B-roll. Talking head vs B-roll covers how the two fit together.
The format fits when the speaker's words are the content: a founder announcing a change, a teacher walking through a lesson, a support lead answering a common question, a creator reviewing a product.
How do you make a talking head video without filming?
Start from what you have. With only a script, an AI avatar can say it. With the speech as an audio file, a still photo or an avatar can be lip-synced to it; on Sume that audio has to be Sume-hosted, such as speech made with TTS 1.0. What doesn't work is a generated video clip with narration laid underneath: Sume's model guide says video models do not lip-sync to generated TTS or to a later voice-over. Talking head video API: avatar vs lip sync vs motion control compares Sume's routes in detail.
| Way to make it | What you need | On Sume |
|---|---|---|
| Film it | A person, a camera, and a quiet room | Not a Sume feature |
| AI avatar from a script | The words and a presenter | An Avatar 1.0 talking video: an estimated 4–60 seconds, English speech |
| Lip sync a still to speech | A photo or an avatar, plus the speech as audio | VEED Fabric 1.0: 1–300 seconds of Sume-hosted audio, such as TTS 1.0 output, at most 10 MB |
What shape should a talking head video be?
It follows where the video plays: vertical for phone feeds, 16:9 for a website or a slide deck. Sume avatar videos default to vertical 9:16 and come in four other shapes; Talking head video format covers the shapes, sizes and files.
What are the limits of an AI talking head on Sume?
An Avatar 1.0 talking head runs an estimated 4–60 seconds with one avatar and one shared scene, speaks English only in current code, and is rendered as a job you poll rather than live. What is an AI avatar? lists these limits, and How to make an AI avatar video longer than 60 seconds joins several parts into one video.
Sources
Related posts
More in Sume Avatar 1.0
- What is AI UGC? Meaning, how it's made, and how it differs
AI UGC is UGC-style video made with AI: a generated presenter speaks your script to camera, the way a creator would on a phone. How it differs.
- What is an AI digital twin? Twins vs photo and stock avatars
An AI digital twin is a custom avatar of one real person that speaks new scripts in their likeness. HeyGen trains it on footage; Synthesia starts from a photo.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume