Split one voiceover into scene clips with timeline audio ranges
Sume timeline audio split slices one hosted file into up to 20 ranges for $0.01 per job, and ranges may overlap. For a video source, detach its audio first.

Call the timeline audio endpoint with operation split, the URL of one Sume-hosted audio file, and up to 20 ranges; each range comes back as its own file. Ranges may overlap, so adjacent scenes can share a short stretch of audio for a crossfade. It costs a flat $0.01 per job.
For a talking-head video, run audio detach once to get the audio, then split it here.
The request
The timeline audio docs require operation: split, a top-level url and ranges of 1 to 20 items. Each range has a start and an optional end; leaving the end out means the rest of the file. Do not send parts, which belongs to concat and fails with audio_split_takes_no_parts. A range whose end is at or before its start returns audio_range_end_before_start. URLs must be Sume-hosted; an off-host URL is rejected at admission, so import it first.
| Input | Rule | Error if wrong |
|---|---|---|
| operation | split | audio_split_requires_url or audio_split_requires_ranges when fields are missing |
| ranges | 1 to 20, each with start and optional end | audio_range_end_before_start when end is not after start |
| parts | Must be absent | audio_split_takes_no_parts |
| url | Sume-hosted audio | unsupported_media_source or source_not_found |
What you get back
The result is a timeline_audio kind with segments[], each segment carrying its own audio_url. Output defaults to wav in pcm_s16le, which is sample-exact. mp3 is available but re-adds priming padding at every edge, so keep wav if the pieces will be joined again or will drive lip-sync. Produced audio is limited to 1800 seconds.
Using overlap on purpose
Because ranges may overlap, you can cut scene one from 0 to 12.4 seconds and scene two from 12.2 seconds onward, giving the editor 200 ms of shared audio for a crossfade. Without overlap, a hard cut on the boundary can clip a breath or a consonant.
- Start from sentence timings, not from visual cuts.
- Overlap by a few hundred milliseconds when you plan a crossfade.
- Leave the last range open-ended to avoid trimming the tail.
- Count your scenes first; the limit is 20 per call.
From a video source
If the voice lives inside an MP4, extract it with audio detach, which limits the source to 1800 seconds and the output to 900, then split the resulting file. The docs recommend exactly that sequence for many ranges off a talking-head video.
Related posts
More in Media tools
- Split one voiceover into six beat files with Timeline audio split
Slice a 180-second voice file into up to 20 ranges for $0.01 flat, then use each file as a beat of a Reel. Ranges can overlap and the last can be open-ended.
- Spotify podcast transcript upload: VTT, 5MB and Sume STT parts
Spotify takes VTT or SRT up to 5MB, with timestamps. Build one from Sume STT in 10-minute parts, stitch the cues, and upload from Spotify for Creators.
- SRT or WebVTT? What Vimeo, Spotify, Apple and Cloudflare accept
Vimeo, Spotify and Apple Podcasts take SRT or WebVTT; Cloudflare Stream documents WebVTT. A table read 2026-10-03 and a script writing both from Sume segments.
- Swap the model in a clothing video with AI: H3 Max Recast
Recast swaps the person in a clip for a person from a photo, 5 to 30 seconds. What it keeps, what the docs leave open about the garment, and the Sume call.
Written by Sume