Split one voiceover into 40 scene tracks: two timeline audio jobs

Timeline audio split takes up to 20 ranges per job. Forty scene tracks from one voiceover is two split jobs, $0.02, with ranges built from cue times in Python.

5 min readSume
All posts

To cut one long voiceover into 40 scene tracks, send two split jobs to Sume's timeline audio endpoint, because each takes 1 to 20 ranges. The cost is $0.01 per job, so $0.02, and the output is durable media.sume.com files with segments[] for each range. Build the ranges from your scene cue times, and let the second job cover cues 21 to 40.

This matters when a long narration is generated in one pass, which suits per-request pricing like Microsoft's $15 and $22 per 1M characters (read 2026-10-03), and then edited into scenes.

Rules of the split

For a talking-head MP4, detach the audio once and split that, rather than detaching 40 times. If the voiceover came from TTS, generate it as one file or a few, then split.

From Sume's Timeline audio and Audio detach docs, read 2026-10-03.
RuleValue
Required fieldsoperation: "split", top-level url, ranges[]
Ranges per job1 to 20
Range shape{ start, end? }; no end means the rest of the file
OverlapAllowed
parts in a splitRejected (audio_split_takes_no_parts)
SourceSume-hosted audio on media.sume.com
Outputwav default, or mp3; produced audio up to 1,800 s
Price$0.01 flat per job

Build the ranges from cue times

This script takes scene start times and produces the request bodies in chunks of 20. The last range of the final chunk leaves end open, so it runs to the end of the file. The first chunk's last range ends at the first cue of the second chunk.

import json

starts = [round(i * 7.5, 2) for i in range(40)]  # scene start times in seconds

def ranges(starts):
    out = []
    for i, s in enumerate(starts):
        r = {"start": s}
        if i + 1 < len(starts):
            r["end"] = starts[i + 1]
        out.append(r)
    return out

all_ranges = ranges(starts)
url = "https://media.sume.com/artifacts/artf_demo/voiceover.wav"
bodies = [
    {"operation": "split", "url": url, "ranges": all_ranges[i:i + 20]}
    for i in range(0, len(all_ranges), 20)
]
print(len(bodies), "jobs, $%.2f" % (0.01 * len(bodies)))
print(json.dumps(bodies[1]["ranges"][-2:]))

Gotchas

Send each body to POST /v1/timeline-1.0/audio with its own Idempotency-Key, then poll GET /v1/jobs/:id/status and read /result. Keep scene order when you merge the two result lists: segments[] are indexed within a job, so scene 21 is index 0 of the second job.

Keep wav. The docs warn that mp3 re-adds priming padding at every edge, which matters if the scene tracks are joined again or drive lip-sync. Also remember that Sume's video models do not lip-sync to generated TTS or to a later voice-over, per the models overview, so scene tracks go under B-roll or a Timeline 1.0 spine, not onto a video model clip.

If a scene is retaken, rerun only the job that contains it. Use overlapping ranges for a small margin around cuts, since the docs allow overlap, and trim in the edit rather than re-splitting.

Checking the split

After both jobs finish, add up the segments[] durations and compare them with the source. With non-overlapping ranges the sum should match the voiceover length to within a frame or so. If you used overlap for margins, the sum will be larger by exactly the overlap you added. A quick script that prints any scene shorter than half a second catches cue lists with duplicate or out-of-order start times.

Name your scene files by scene number as soon as you receive them, and store the job id with each. If a client asks for scene 17 to be re-cut, you will want to find the one job to rerun, which costs a cent.

Alternative: split inside the render

If the scene audio is only needed for one final video, check whether Timeline 1.0 can take the voiceover as its audio spine with the video slots timed against it, as the Timeline docs describe, instead of producing 40 files you will not reuse. Separate scene files make sense when different tools or people take the scenes, or when you want to regenerate single scenes with new takes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume