Split one voiceover into six beat files with Timeline audio split

Slice a 180-second voice file into up to 20 ranges for $0.01 flat, then use each file as a beat of a Reel. Ranges can overlap and the last can be open-ended.

5 min readSume
All posts

Sume's timeline audio split cuts one hosted audio file into up to 20 ranges in a single job and costs a flat $0.01. Each range returns its own durable audio URL, so a 3-minute voiceover can become six beat files in one call. Ranges may overlap, and a range with no end runs to the end of the file.

Why split a voice file

A Reel at 3 minutes (reported by Metricool (read 2026-10-03)) is easier to edit as beats, where each beat's pictures start where its voice starts. If you recorded one long take or detached one from an earlier Reel, splitting gives you per-beat files for avatar jobs, captions or re-voicing, without running text-to-speech again.

The request

POST /v1/timeline-1.0/audio with operation: "split", a top-level url and ranges[]. Do not send parts, which belongs to concat. An end at or before the start fails with audio_range_end_before_start. The default output is wav; mp3 is available through output.format but is smaller and less exact.

Six ranges for a 180-second voice file (example values)
Beatstart (s)end (s)
1022
22252
35288
488124
5124158
6158(omit)

Code

Idempotency is required, so a repeated call with the same key does not bill twice.

import os, requests

edges = [0, 22, 52, 88, 124, 158]
ranges = []
for i, start in enumerate(edges):
    r = {"start": start}
    if i + 1 < len(edges):
        r["end"] = edges[i + 1]
    ranges.append(r)

res = requests.post(
    "https://api.sume.com/v1/timeline-1.0/audio",
    headers={
        "Authorization": "Bearer " + os.environ["SUME_API_KEY"],
        "Idempotency-Key": "split-beats-001",
    },
    json={
        "operation": "split",
        "url": os.environ["VOICE_URL"],
        "ranges": ranges,
    },
)
res.raise_for_status()
print(res.json())

What Sume does not do

Split does not find the beat boundaries. You choose them, for example from the sentence segments of a transcript. It also does not re-time: the produced audio is limited to 1800 seconds, and the cut points are exactly the ones you send. For a join that is only needed inside one render, skip the audio job and use audio.parts[] in Timeline 1.0 (up to 20 slices).

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume