What are faceless videos? The formats and how each is made
Faceless videos tell a story without the creator on camera: voiceover over B-roll, on-screen text, illustrated scenes, or an AI presenter.

Faceless videos are videos that carry their message without the creator appearing on camera. The story is told by a voiceover over B-roll footage, by text and captions on screen, by illustrated or animated scenes, by a screen recording, or by a synthetic presenter standing in for the creator.
Each format decides what you have to produce: a voice, pictures that match the words, captions, and usually a music bed. The Sume facts below come from the Video generation, Video captions, Timeline 1.0, and Generate avatar video docs and the Sume API reference, read on 2026-09-28.
What does faceless mean for a channel?
A faceless channel publishes videos where no face identifies the creator: the viewer hears a narrator and sees footage, text, or illustrations. It lets a creator stay private, publish without filming themselves, and show a topic such as history, finance facts, or stories instead of presenting it.
Faceless does not mean silent. Most faceless formats lean on a voice, and the pictures follow what the voice says. That is why the same few ingredients come back in every format.
What are the main faceless video formats?
These five formats are common. The last column is the Sume call that produces each part; a screen recording is the one format you capture yourself.
| Format | On screen | What you need | Sume call |
|---|---|---|---|
| Voiceover over B-roll | Footage that illustrates the narration | A voice, a clip per line or beat, a music bed | POST /v1/tts-1.0/generate, POST /v1/videos, POST /v1/timeline-1.0/render |
| Text and captions | Words on screen over a simple background | The copy with a start and end time for each line | POST /v1/video-captions with cues |
| Illustrated or animated scenes | A generated still per scene, animated | A still per scene, then a clip that opens on it | POST /v1/images, then POST /v1/videos with a first_frame |
| Synthetic presenter | An AI avatar speaking the script | A ready avatar and a script | POST /v1/avatar-1.0/talking-video |
| Screen recording | Your screen, with a voiceover | Your own capture | None: Sume does not record screens |
How is each part made with AI?
The voice, the pictures, the captions, and the assembly are separate steps, and each is a separate job on Sume. Automate faceless short-form videos walks through every call, its fields, and its cost; in short:
- Voice: text to speech reads the script. Sume's TTS 1.0 can borrow an avatar's voice, and only the voice: no face appears.
- Pictures: a video model generates a clip per line from a prompt, in
9:16for vertical. A still sent as the clip'sfirst_framemakes it open on that picture. - Captions: burned in from the speech, or from your own words and timings exactly as you send them.
- Assembly: one render lays the clips over the voice with a music bed that dips while the voice speaks (What is audio ducking?). In current code the render's sound is the voice plus the bed; each clip's own audio is dropped.
Is an AI avatar video still faceless?
It depends on what you mean. The creator stays off camera, but a face is on screen. Sume's Avatar 1.0 turns a ready avatar and a script into a talking video when Sume estimates the result at 4–60 seconds, so a longer script is split into several jobs. In current code the avatar speaks English only. Talking head vs B-roll explains how a presenter shot and B-roll fit together.
What are some faceless video ideas?
Pick a subject where the pictures can follow the words. Each of these has its own guide:
- History facts and period stories: How to make AI history videos.
- Long narrated explainers: How to make an AI documentary video.
- Quotes and lessons over cinematic shots: Create motivational videos with AI.
- A daily Short on one topic, run on a schedule: Automate faceless short-form videos.
What are the limits?
- In current code a caption job refuses a video longer than 60 seconds or one with no audio stream, so a text-only video needs a sound track before captions.
- Timeline 1.0 reads only files already on Sume, such as the output of an earlier Sume job.
- Generated footage illustrates the words; it is not a record of real events.
Sources
Related posts
More in Use cases
- What is a hero video? The website loop and the brand film
A hero video is the short, usually silent, looping clip at the top of a web page. In campaign planning, the hero video is the flagship film instead.
- What is a packshot? Meaning, examples, and use in AI
A packshot is a clean, evenly lit photo of a product, often in its packaging, on a plain background. AI tools use one as the product's reference.
- What is creative automation? Templates, data, and review
Creative automation makes marketing assets by running one template over changing data, with people approving results instead of building each one.
- What is virtual try-on? AR try-on and AI try-on explained
Virtual try-on shows a product on a person who isn't wearing it: live on a camera feed with AR, or in a new AI-generated image or video made from photos.
Written by Sume