Talking head vs B-roll: what's the difference?
A talking head is the shot of someone speaking to camera; B-roll is the footage you cut to while their voice keeps playing. How the two fit together.

A talking head is the shot of a person speaking to the camera; B-roll is the supporting footage you cut to while that person's voice keeps playing. The talking head, also called A-roll, carries the words. B-roll shows what the words are about: the product, the place, the steps.
An AI talking-head video follows the same rule: the avatar's voice runs as one continuous track, and the picture cuts between the avatar and the B-roll over it, so a cutaway never cuts the words. The Sume facts come from the Generate avatar video, Timeline 1.0, Audio detach and Models overview docs, read on 2026-09-28; the definitions are general. Anything called current behavior is read from Sume's code.
What is the difference between a talking head and B-roll?
One is the person saying it, the other is the thing being said. In the edit, the voice stays constant and the picture switches between them.
| Talking head (A-roll) | B-roll | |
|---|---|---|
| On screen | The speaker, facing the camera | The product, the place, the steps, a detail |
| Sound | The voice the viewer follows | The speaker's voice keeps playing over it |
| Job in the edit | Carries the words and the presence | Shows what the words describe and hides the cuts |
| Made with Sume | An avatar talking video, or a still lip-synced with VEED Fabric 1.0 | Image and video generation: wordless shots need no lip sync |
Is a talking head the same as A-roll?
In a presenter or interview video, yes. A-roll is the primary footage, the shots the story is told through, and there that is the person talking. B-roll is everything you cut to. In a video with no speaker on screen, such as narration over footage, there is no talking head at all; a faceless video is built from voiceover and B-roll alone.
How do a talking head and B-roll work together?
The voice runs without a break, and the picture cuts away and comes back. When it comes back to the speaker, the lips have to match the words at that moment, so the speaker's clip resumes at the same point in time as the voice. If the presenter should stay visible during a cutaway, put both in one frame, picture in picture or split screen.
- Stay on the talking head for the opening, a direct question to the viewer, and the call to action.
- Cut to B-roll when the words describe something you can show: a feature, a step, a place, a number.
- Use B-roll to cover a join, for example where two separately rendered clips meet in a longer video.
How do I add B-roll to an AI talking-head video?
A Sume avatar video resolves one avatar and one shared scene, so anything else on screen, B-roll included, is cut in afterwards. The edit is one Timeline 1.0 render with the avatar's voice, detached from its clip with audio detach, as the audio spine, and every clip in it must already be one of your workspace's media.sume.com files, such as B-roll generated on Sume. Keep the voice on the spine: in current code a render plays only the spine and an optional soundtrack, and each clip's own audio is dropped. Add B-roll to a talking-head avatar video has the slot layout and the cut-back timing.
Sources
Related posts
More in Sume Avatar 1.0
- UGC video prompt for AI avatars: what goes in each field
A UGC video prompt for an AI avatar is three inputs, not one: the creator's words, a scene prompt for the setting and light, and a product image.
- What is AI UGC? Meaning, how it's made, and how it differs
AI UGC is UGC-style video made with AI: a generated presenter speaks your script to camera, the way a creator would on a phone. How it differs.
- What is an AI avatar? What it does and what it can't
An AI avatar is a digital presenter, a face and a voice, that speaks the script you write in a video. What it does, how it's made, and its limits.
- What is an AI digital twin? Twins vs photo and stock avatars
An AI digital twin is a custom avatar of one real person that speaks new scripts in their likeness. HeyGen trains it on footage; Synthesia starts from a photo.
Written by Sume