Split one voiceover into 40 scene tracks: two timeline audio jobs
Timeline audio split takes up to 20 ranges per job. Forty scene tracks from one voiceover is two split jobs, $0.02, with ranges built from cue times in Python.

To cut one long voiceover into 40 scene tracks, send two split jobs to Sume's timeline audio endpoint, because each takes 1 to 20 ranges. The cost is $0.01 per job, so $0.02, and the output is durable media.sume.com files with segments[] for each range. Build the ranges from your scene cue times, and let the second job cover cues 21 to 40.
This matters when a long narration is generated in one pass, which suits per-request pricing like Microsoft's $15 and $22 per 1M characters (read 2026-10-03), and then edited into scenes.
Rules of the split
For a talking-head MP4, detach the audio once and split that, rather than detaching 40 times. If the voiceover came from TTS, generate it as one file or a few, then split.
| Rule | Value |
|---|---|
| Required fields | operation: "split", top-level url, ranges[] |
| Ranges per job | 1 to 20 |
| Range shape | { start, end? }; no end means the rest of the file |
| Overlap | Allowed |
parts in a split | Rejected (audio_split_takes_no_parts) |
| Source | Sume-hosted audio on media.sume.com |
| Output | wav default, or mp3; produced audio up to 1,800 s |
| Price | $0.01 flat per job |
Build the ranges from cue times
This script takes scene start times and produces the request bodies in chunks of 20. The last range of the final chunk leaves end open, so it runs to the end of the file. The first chunk's last range ends at the first cue of the second chunk.
import json
starts = [round(i * 7.5, 2) for i in range(40)] # scene start times in seconds
def ranges(starts):
out = []
for i, s in enumerate(starts):
r = {"start": s}
if i + 1 < len(starts):
r["end"] = starts[i + 1]
out.append(r)
return out
all_ranges = ranges(starts)
url = "https://media.sume.com/artifacts/artf_demo/voiceover.wav"
bodies = [
{"operation": "split", "url": url, "ranges": all_ranges[i:i + 20]}
for i in range(0, len(all_ranges), 20)
]
print(len(bodies), "jobs, $%.2f" % (0.01 * len(bodies)))
print(json.dumps(bodies[1]["ranges"][-2:]))
Gotchas
Send each body to POST /v1/timeline-1.0/audio with its own Idempotency-Key, then poll GET /v1/jobs/:id/status and read /result. Keep scene order when you merge the two result lists: segments[] are indexed within a job, so scene 21 is index 0 of the second job.
Keep wav. The docs warn that mp3 re-adds priming padding at every edge, which matters if the scene tracks are joined again or drive lip-sync. Also remember that Sume's video models do not lip-sync to generated TTS or to a later voice-over, per the models overview, so scene tracks go under B-roll or a Timeline 1.0 spine, not onto a video model clip.
If a scene is retaken, rerun only the job that contains it. Use overlapping ranges for a small margin around cuts, since the docs allow overlap, and trim in the edit rather than re-splitting.
Checking the split
After both jobs finish, add up the segments[] durations and compare them with the source. With non-overlapping ranges the sum should match the voiceover length to within a frame or so. If you used overlap for margins, the sum will be larger by exactly the overlap you added. A quick script that prints any scene shorter than half a second catches cue lists with duplicate or out-of-order start times.
Name your scene files by scene number as soon as you receive them, and store the job id with each. If a client asks for scene 17 to be re-cut, you will want to find the one job to rerun, which costs a cent.
Alternative: split inside the render
If the scene audio is only needed for one final video, check whether Timeline 1.0 can take the voiceover as its audio spine with the video slots timed against it, as the Timeline docs describe, instead of producing 40 files you will not reuse. Separate scene files make sense when different tools or people take the scenes, or when you want to regenerate single scenes with new takes.
Sources
Related posts
More in Developers
- Spring AI MCP client request-timeout 20s vs Sume jobs_wait
Spring AI's MCP client defaults request-timeout to 20s, shorter than a 50s Sume jobs_wait. Raise it with a customizer or keep waits short and re-issue them.
- SSML in text to speech: Sume takes a plain transcript, no ssml field
Does Sume's text to speech accept SSML? The tts_create body has a plain transcript and rejects unknown keys. What to use for speed, volume, emotion and pauses.
- Stippled AI graphics turn to gray mush when resized: a downscale test
A dotted AI graphic can lose most of its contrast when downscaled. A Pillow test of nearest, bilinear and Lanczos on a stipple, and what to request instead.
- Streaming transcripts: burn only final text into clip captions
MAI-Transcribe-2-Streaming returns partials in about 100 ms, then settled text. Burned-in captions need the final text. Filter it and pass it to Sume as cues.
Written by Sume