What are faceless videos? The formats and how each is made

Faceless videos tell a story without the creator on camera: voiceover over B-roll, on-screen text, illustrated scenes, or an AI presenter.

4 min readSume
All posts

Faceless videos are videos that carry their message without the creator appearing on camera. The story is told by a voiceover over B-roll footage, by text and captions on screen, by illustrated or animated scenes, by a screen recording, or by a synthetic presenter standing in for the creator.

Each format decides what you have to produce: a voice, pictures that match the words, captions, and usually a music bed. The Sume facts below come from the Video generation, Video captions, Timeline 1.0, and Generate avatar video docs and the Sume API reference, read on 2026-09-28.

What does faceless mean for a channel?

A faceless channel publishes videos where no face identifies the creator: the viewer hears a narrator and sees footage, text, or illustrations. It lets a creator stay private, publish without filming themselves, and show a topic such as history, finance facts, or stories instead of presenting it.

Faceless does not mean silent. Most faceless formats lean on a voice, and the pictures follow what the voice says. That is why the same few ingredients come back in every format.

What are the main faceless video formats?

These five formats are common. The last column is the Sume call that produces each part; a screen recording is the one format you capture yourself.

Sume calls from Video generation, Video captions, Timeline 1.0, Generate avatar video, and the Sume API reference, read 2026-09-28.
FormatOn screenWhat you needSume call
Voiceover over B-rollFootage that illustrates the narrationA voice, a clip per line or beat, a music bedPOST /v1/tts-1.0/generate, POST /v1/videos, POST /v1/timeline-1.0/render
Text and captionsWords on screen over a simple backgroundThe copy with a start and end time for each linePOST /v1/video-captions with cues
Illustrated or animated scenesA generated still per scene, animatedA still per scene, then a clip that opens on itPOST /v1/images, then POST /v1/videos with a first_frame
Synthetic presenterAn AI avatar speaking the scriptA ready avatar and a scriptPOST /v1/avatar-1.0/talking-video
Screen recordingYour screen, with a voiceoverYour own captureNone: Sume does not record screens

How is each part made with AI?

The voice, the pictures, the captions, and the assembly are separate steps, and each is a separate job on Sume. Automate faceless short-form videos walks through every call, its fields, and its cost; in short:

  • Voice: text to speech reads the script. Sume's TTS 1.0 can borrow an avatar's voice, and only the voice: no face appears.
  • Pictures: a video model generates a clip per line from a prompt, in 9:16 for vertical. A still sent as the clip's first_frame makes it open on that picture.
  • Captions: burned in from the speech, or from your own words and timings exactly as you send them.
  • Assembly: one render lays the clips over the voice with a music bed that dips while the voice speaks (What is audio ducking?). In current code the render's sound is the voice plus the bed; each clip's own audio is dropped.

Is an AI avatar video still faceless?

It depends on what you mean. The creator stays off camera, but a face is on screen. Sume's Avatar 1.0 turns a ready avatar and a script into a talking video when Sume estimates the result at 4–60 seconds, so a longer script is split into several jobs. In current code the avatar speaks English only. Talking head vs B-roll explains how a presenter shot and B-roll fit together.

What are some faceless video ideas?

Pick a subject where the pictures can follow the words. Each of these has its own guide:

What are the limits?

  • In current code a caption job refuses a video longer than 60 seconds or one with no audio stream, so a text-only video needs a sound track before captions.
  • Timeline 1.0 reads only files already on Sume, such as the output of an earlier Sume job.
  • Generated footage illustrates the words; it is not a record of real events.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume