Voiceover longer than your clips: Timeline pads, loops, 0.5 s rule
If the voiceover outruns your clips, Sume Timeline still renders to audio.duration_seconds: short sources pad or loop, and coverage may stop 0.5 s early.

When a voiceover is longer than the clips you cut under it, a Sume Timeline 1.0 render still runs for the length you declare in audio.duration_seconds. Short sources are padded or looped and reported as soft warnings, not failures. But the video[] slots themselves can stop at most 0.5 seconds before the end of the voiceover, so you have to plan slot lengths that reach the end.
This follows the Timeline 1.0 docs (read 2026-10-06).
What decides the length of the output?
audio.duration_seconds is the output length. It is required and runs 1 to 1,800 seconds. Take it from the TTS result: with timestamps.words on, the last word's end gives the spoken length, and the result also reports duration_seconds. The render is priced at $0.10 per ceil output minute, and the reserve is ceil(audio.duration_seconds / 60) minutes.
What happens to a short clip?
The docs say soft warnings include padded or looped short sources, snapped transitions and ignored still motion. They are not failures. The render finishes, so check warnings[] in the result and look at the output before you publish.
| Item | Rule | What to do |
|---|---|---|
video[0].start | Must be 0 | Start the first slot at the top of the voiceover |
Later video[].start | Must increase | Declared starts are authoritative |
| Coverage at the end | May stop at most 0.5 s before the end of the spine | Add a slot or lengthen the last duration |
| Short source | Padded or looped, with a warning | Read warnings[] and inspect the cut |
How do I plan the slots?
Add up the slot durations first. Timeline has a plan step that checks URLs and compiles the program without creating a job or reserving credits; it returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros. Run it before the paid render. If you need to join several voiceover lines into one spine first, Timeline audio docs cover a gapless concat that returns segments[] offsets you can use to place the clips.
Sources
More in Media tools
- Why a voiceover on an AI video clip never lip-syncs
Video models do not lip-sync to TTS or a later voiceover. Per Sume's docs a talking face comes from Fabric, H3 Max Lip Sync or an avatar talking-video job.
- Public URL or media.sume.com: which Sume media tool takes which
Sume video trim and timeline accept only media.sume.com URLs from your workspace. Captions and video upscale take a public HTTPS URL. Import first.
- Why every AI background track sounds the same: a seven-axis brief
If your AI music all sounds alike, the prompt is probably generic. Sume's Lyria brief has seven axes: emotion, genre, BPM, key, instruments, arc and era.
- Write a YouTube show description from your episode transcripts
YouTube asks you to set a show's details but lists no required fields. Pull transcripts with Sume video inspect at $0.01 a minute and draft from them.
Written by Sume