3-minute explainer Reel as six beats: voice first, then pictures

Plan a 180-second Reel as six voiced beats: one TTS file per beat, one concat for offsets, then picture slots that start where each beat starts.

5 min readSume
All posts

The reliable way to build a 3-minute explainer Reel is to make the voice first and let it set the clock: write six beats of about 30 seconds each, generate one voice file per beat, join them with Sume's timeline audio concat, and place each beat's pictures at the offset the concat returns. Offsets from the audio, not guesses, keep the pictures on the words.

Instagram's Reels guide (read 2026-10-03) says 60 seconds to 3 minutes suits longer storytelling, walkthroughs and detailed how-tos, and Metricool's Instagram news (read 2026-10-03) reports the 3-minute length. A beat structure gives a long Reel the rhythm of several short ones.

A beat sheet that fits one concat

Timeline audio concat accepts 1 to 20 ordered parts, so six beats leave plenty of room, including for a short breath file between beats. Each part is one TTS job, which keeps every file well under the 20,000-character request cap and lets you redo one beat without touching the rest.

Six-beat plan for a 180-second Reel (planning values, read 2026-10-03)
BeatJob in the storyTarget lengthPicture idea
1Hook: the problem in one lineabout 20 sFast cuts, 3 to 4 slots
2Why it happensabout 30 sTwo or three explainer shots
3The fix, step oneabout 35 sScreen or product close-up
4The fix, step twoabout 35 sSecond angle, same look
5Proof or resultabout 35 sBefore and after
6What to do nextabout 25 sEnd card hold

Join the voices and read the offsets

POST /v1/timeline-1.0/audio with operation: "concat" joins parts in the sample domain, with no re-synthesis and no silence at the seams, for a flat $0.01. The result carries one audio_url, a duration_seconds, and segments[] with each part's index, start and duration_seconds. Those start values are the numbers you re-base the picture slots against. Parts must share one channel layout, or the job fails audio_parts_channel_mismatch; ask for the same output_format on every TTS job so they do. Produced audio can be up to 1800 seconds.

def slots(segments, scenes_per_beat):
    """segments: concat result. scenes_per_beat: list of lists of urls."""
    video = []
    for seg, urls in zip(segments, scenes_per_beat):
        step = seg["duration_seconds"] / len(urls)
        for i, url in enumerate(urls):
            video.append({
                "source_url": url,
                "start": round(seg["start"] + i * step, 3),
                "duration": round(step, 3),
            })
    return video

segs = [{"index": 0, "start": 0, "duration_seconds": 21.4},
        {"index": 1, "start": 21.4, "duration_seconds": 30.2}]
scenes = [["https://media.sume.com/artifacts/artf_demo/a.mp4"],
          ["https://media.sume.com/artifacts/artf_demo/b.mp4",
           "https://media.sume.com/artifacts/artf_demo/c.mp4"]]
print(slots(segs, scenes))

Render and check the length

Feed the joined audio_url into Timeline 1.0 as audio.url with audio.duration_seconds set from the concat result. The picture slots have to cover the spine: coverage may trail it by at most 0.5 seconds, so let the last slot run to the end. Timeline bills $0.10 per output minute rounded up, so a 178-second file costs $0.30, the same as exactly 180. Read the warnings[] on the result for padded or looped short sources, and poll the job as described in Jobs and results.

What Sume does and does not do

Sume gives you the voice jobs, the gapless join, the offsets and the render. It does not decide where a beat should end, lip-sync b-roll to the voice, or shorten a beat that ran long. When a beat overruns, rewrite its script or nudge generation_config.speed (0.6 to 1.5) and regenerate that one file; the concat and the render are cheap to repeat.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume