3-minute explainer Reel as six beats: voice first, then pictures
Plan a 180-second Reel as six voiced beats: one TTS file per beat, one concat for offsets, then picture slots that start where each beat starts.

The reliable way to build a 3-minute explainer Reel is to make the voice first and let it set the clock: write six beats of about 30 seconds each, generate one voice file per beat, join them with Sume's timeline audio concat, and place each beat's pictures at the offset the concat returns. Offsets from the audio, not guesses, keep the pictures on the words.
Instagram's Reels guide (read 2026-10-03) says 60 seconds to 3 minutes suits longer storytelling, walkthroughs and detailed how-tos, and Metricool's Instagram news (read 2026-10-03) reports the 3-minute length. A beat structure gives a long Reel the rhythm of several short ones.
A beat sheet that fits one concat
Timeline audio concat accepts 1 to 20 ordered parts, so six beats leave plenty of room, including for a short breath file between beats. Each part is one TTS job, which keeps every file well under the 20,000-character request cap and lets you redo one beat without touching the rest.
| Beat | Job in the story | Target length | Picture idea |
|---|---|---|---|
| 1 | Hook: the problem in one line | about 20 s | Fast cuts, 3 to 4 slots |
| 2 | Why it happens | about 30 s | Two or three explainer shots |
| 3 | The fix, step one | about 35 s | Screen or product close-up |
| 4 | The fix, step two | about 35 s | Second angle, same look |
| 5 | Proof or result | about 35 s | Before and after |
| 6 | What to do next | about 25 s | End card hold |
Join the voices and read the offsets
POST /v1/timeline-1.0/audio with operation: "concat" joins parts in the sample domain, with no re-synthesis and no silence at the seams, for a flat $0.01. The result carries one audio_url, a duration_seconds, and segments[] with each part's index, start and duration_seconds. Those start values are the numbers you re-base the picture slots against. Parts must share one channel layout, or the job fails audio_parts_channel_mismatch; ask for the same output_format on every TTS job so they do. Produced audio can be up to 1800 seconds.
def slots(segments, scenes_per_beat):
"""segments: concat result. scenes_per_beat: list of lists of urls."""
video = []
for seg, urls in zip(segments, scenes_per_beat):
step = seg["duration_seconds"] / len(urls)
for i, url in enumerate(urls):
video.append({
"source_url": url,
"start": round(seg["start"] + i * step, 3),
"duration": round(step, 3),
})
return video
segs = [{"index": 0, "start": 0, "duration_seconds": 21.4},
{"index": 1, "start": 21.4, "duration_seconds": 30.2}]
scenes = [["https://media.sume.com/artifacts/artf_demo/a.mp4"],
["https://media.sume.com/artifacts/artf_demo/b.mp4",
"https://media.sume.com/artifacts/artf_demo/c.mp4"]]
print(slots(segs, scenes))Render and check the length
Feed the joined audio_url into Timeline 1.0 as audio.url with audio.duration_seconds set from the concat result. The picture slots have to cover the spine: coverage may trail it by at most 0.5 seconds, so let the last slot run to the end. Timeline bills $0.10 per output minute rounded up, so a 178-second file costs $0.30, the same as exactly 180. Read the warnings[] on the result for padded or looped short sources, and poll the job as described in Jobs and results.
What Sume does and does not do
Sume gives you the voice jobs, the gapless join, the offsets and the render. It does not decide where a beat should end, lip-sync b-roll to the voice, or shorten a beat that ran long. When a beat overruns, rewrite its script or nudge generation_config.speed (0.6 to 1.5) and regenerate that one file; the concat and the render are cheap to repeat.
Sources
Related posts
More in Use cases
- 3-minute Reel script with AI voice: characters, cost and the cap
A 180-second Reel narration is a character budget, not a word count. Sume TTS is $0.0475 per 1,000 characters with a 20,000-character cap; measure, then scale.
- A/B test three Shorts hooks from one body: Timeline source_in
YouTube announced A/B testing for up to three Short versions. Build the three files from one body with Timeline slots, source_in and a plan call before you pay.
- Add a related video to a YouTube Short made outside YouTube
A related video is a clickable link under a Short's channel name. YouTube says it needs advanced features and a public or unlisted video. Set it in Studio.
- Add AI voiceover in another language to a silent screen recording
Write the narration, generate it with a language code, lay it over the recording in a Timeline render, then caption it. What each step bills and what it limits.
Written by Sume