Split one voiceover into scene clips with timeline audio ranges

Sume timeline audio split slices one hosted file into up to 20 ranges for $0.01 per job, and ranges may overlap. For a video source, detach its audio first.

4 min readSume
All posts

Call the timeline audio endpoint with operation split, the URL of one Sume-hosted audio file, and up to 20 ranges; each range comes back as its own file. Ranges may overlap, so adjacent scenes can share a short stretch of audio for a crossfade. It costs a flat $0.01 per job.

For a talking-head video, run audio detach once to get the audio, then split it here.

The request

The timeline audio docs require operation: split, a top-level url and ranges of 1 to 20 items. Each range has a start and an optional end; leaving the end out means the rest of the file. Do not send parts, which belongs to concat and fails with audio_split_takes_no_parts. A range whose end is at or before its start returns audio_range_end_before_start. URLs must be Sume-hosted; an off-host URL is rejected at admission, so import it first.

Split inputs and the errors they raise (read 2026-10-03)
InputRuleError if wrong
operationsplitaudio_split_requires_url or audio_split_requires_ranges when fields are missing
ranges1 to 20, each with start and optional endaudio_range_end_before_start when end is not after start
partsMust be absentaudio_split_takes_no_parts
urlSume-hosted audiounsupported_media_source or source_not_found

What you get back

The result is a timeline_audio kind with segments[], each segment carrying its own audio_url. Output defaults to wav in pcm_s16le, which is sample-exact. mp3 is available but re-adds priming padding at every edge, so keep wav if the pieces will be joined again or will drive lip-sync. Produced audio is limited to 1800 seconds.

Using overlap on purpose

Because ranges may overlap, you can cut scene one from 0 to 12.4 seconds and scene two from 12.2 seconds onward, giving the editor 200 ms of shared audio for a crossfade. Without overlap, a hard cut on the boundary can clip a breath or a consonant.

  • Start from sentence timings, not from visual cuts.
  • Overlap by a few hundred milliseconds when you plan a crossfade.
  • Leave the last range open-ended to avoid trimming the tail.
  • Count your scenes first; the limit is 20 per call.

From a video source

If the voice lives inside an MP4, extract it with audio detach, which limits the source to 1800 seconds and the output to 900, then split the resulting file. The docs recommend exactly that sequence for many ranges off a talking-head video.

Related posts

More in Media tools

All Media tools posts

Written by Sume