Voiceover-only Short: silent B-roll plus a TTS spine in Timeline
A Short with no on-camera speech is narration over B-roll. In Timeline the narration is the audio spine, the clips are slots, and audio sets the length.

A voiceover-only Short is easiest to build narration first: generate the speech, measure it, then lay clips over it. In Sume the narration file becomes audio.url with duration_seconds set to its real length, and every clip is a video slot, so picture follows voice and cannot run past it.
YouTube's editor accepts voiceover in its timeline per Enhance your Shorts (read 2026-10-03). The render rules are from the Timeline 1.0 doc.
What order should I work in?
Voice first, because it fixes the length. Then pick clip count from that length, and send the plan call before you pay. The steps:
- Write the script and generate the narration; note its exact duration.
- Cut the slots: for a 24-second voice and six beats, about 4 seconds a slot, each at least 0.2.
- Use
fit: coverfor clips that are already vertical andblurfor landscape clips. - Call
/v1/timeline-1.0/plan(unbilled) and readbillable_minutes. - Render with an Idempotency-Key.
What if the clips are shorter than the voice?
The Timeline doc says video coverage may trail the narration by at most 0.5 seconds, so plan slots that add up to the voice length. A gap between two slots is filled by holding the previous frame and returns timeline_gap_filled; when the clips end before the audio, the last frame is held and the result carries video_coverage_shorter_than_audio. Both are soft warnings in the compiler, not failures. Read warnings[] on the result instead of assuming the picture covers the narration.
How much does it cost?
Timeline is $0.10 per ceil output minute, so a Short up to 60 seconds is $0.10 for the render. The narration and the clips are separate lines: check the live catalog for the speech model you use and for each video model, because they change.
| Step | Endpoint | Billing |
|---|---|---|
| Narration | Speech model in the catalog, or your own recording | Per the catalog price, or free if recorded |
| B-roll clips | Video Router | Per the model's catalog price |
| Assembly | POST /v1/timeline-1.0/render | $0.10 per ceil output minute |
What Sume does not do
Sume does not choose clips from a script or sync cuts to words on its own. You decide each slot's start and duration.
Sources
Related posts
More in Use cases
- Walmart Recognized Reviewer: under 15 reviews, 70% content score
Walmart's Recognized Reviewer now covers items under 15 reviews, if the content quality score is 70% or higher. What to fix first, and what Sume can make.
- Weekly AI release-notes video: 52 x 30 seconds, annual cost on Sume
A weekly 30-second AI presenter release-notes video for a year is 1,560 seconds: $382.20 on Plus, $287.04 on Standard, with a $6.50 music bed and $0.95 avatar.
- A weekly host video while Tavus Griffin is preview-only: Sume, priced
Tavus describes Griffin as a preview with no date or price. For a weekly on-camera update today, Sume Avatar 1.0 costs $3.68 to $11.00 for 20 seconds.
- About this ad: where Meta shows AI info for third-party AI ads
Meta's June 2026 update puts AI info in a new About this ad menu on every ad. What it says about ads made with outside AI tools, in Meta's own words.
Written by Sume