Split a 9-minute recording into three Reels under 3 minutes

Instagram's Reels page says Reels over 3 minutes won't be recommended to new audiences. Find sentence ends with video-inspect STT, then cut with video-trim.

5 min readSume
All posts

Cut the recording into pieces of at most 180 seconds, ending each on a sentence. Instagram's Reels page says Reels over 3 minutes won't be recommended to new audiences, and that you can record one or more clips adding up to 20 minutes (read 2026-10-07). With Sume: one video-inspect call with sentence segmentation (about $0.09 of STT for 9 minutes at the docs' $0.01 per minute), then three video-trim jobs at $0.02.

The cuts are chosen by a few lines of Python from the sentence timings, so no piece ends mid-word.

What Instagram's page says

From Instagram's Reels page. Note that the same page's overview mentions multi-clip videos up to 3 minutes; the other post in this series covers that mismatch.

Instagram Reels lengths on the features page (read 2026-10-07)
StatementValue
Record one or more clips adding up to20 minutes
Not recommended to new audiences beyond3 minutes
Overview linemulti-clip videos up to 3 minutes

Get sentence timings

Send transcribe: true, segmentation: {"mode": "sentence"}, frames: false, and a duration_seconds hint of up to 600 (without it Sume reserves one minute). The response carries transcript.segments[] with start and end and no gaps. A source of 9 minutes is 540 seconds, under the 1,800-second cap.

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: reel-split-inspect-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk-9min.mp4",
    "frames": false,
    "transcribe": true,
    "duration_seconds": 540,
    "segmentation": { "mode": "sentence" }
  }'

Pick cut points

Walk the segments and, for each window, end on the last sentence that finishes within 180 seconds of the window start. The stand-in segments below are invented so the script runs; replace them with transcript.segments.

def windows(segments, limit=180.0):
    out, start, i = [], 0.0, 0
    while i < len(segments):
        end = None
        while i < len(segments) and segments[i]["end"] - start <= limit:
            end = segments[i]["end"]
            i += 1
        if end is None:
            end = segments[i]["end"]
            i += 1
        out.append({"start": start, "duration": round(end - start, 2)})
        start = end
    return out


lengths = [35, 36, 47, 34, 24] * 5
segs, t = [], 0.0
for n in lengths:
    segs.append({"start": t, "end": t + n})
    t += n
for w in windows(segs):
    print(w)

Cut and check

Send each window as a video trim job with start and duration, each with its own idempotency key. Collect results as in Jobs and results, and probe each output to confirm duration_seconds is 180 or less. The stand-in data only illustrates the logic, and your real output depends on your speech.

Titles and openings for each part

Each part starts mid-conversation, so the first sentence of parts two and three may lean on context a new viewer lacks. Read the first sentence of each window from the transcript and, where it begins with "and" or "so", move start to the next sentence.

Captions help here. A caption job on a 60-second-or-shorter clip is $0.20 in the docs, so a 180-second part is three caption jobs if you split it first; check the video-captions page for how the price is stated before you budget.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume