Split a voice track into 20 files for 1 cent: timeline audio split

Sume timeline audio split cuts one audio file into up to 20 ranges in a single $0.01 job, with wav output by default. 20 separate jobs would cost $0.20.

4 min readSume
All posts

To cut one audio file into many pieces, call POST /v1/timeline-1.0/audio with operation: "split", a top-level url, and ranges[] of up to 20 {start, end} entries. Sume returns one audio file per range for a flat $0.01 per job. Twenty ranges in one job cost $0.01; twenty one-range jobs would cost $0.20.

Split request

The audio must already be a media.sume.com file in your workspace. Ranges may overlap, and end can be left out to run to the end of the file. Send no parts (audio_split_takes_no_parts), and make sure each end is after its start (audio_range_end_before_start).

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: split-001" \
  -d '{
    "operation": "split",
    "url": "https://media.sume.com/artifacts/artf_demo/spine.wav",
    "ranges": [{ "start": 0, "end": 12.4 }, { "start": 12.4 }]
  }'

From a video

For a talking-head MP4, the docs say to detach the audio once and then split it. The cost is two jobs: 1 x $0.01 for audio detach plus 1 x $0.01 for the split, $0.02 for 20 files. Detach caps output at 900 seconds, so a track longer than that needs a range on the detach.

Output and cost

The result is kind: timeline_audio with segments[], each with its own audio_url, plus timings. output.format is wav (default, sample-exact PCM) or mp3, which is smaller but adds priming padding at every edge. Keep wav if the pieces will be joined again or will drive lip-sync. The produced audio is at most 1,800 seconds.

Cost of 20 pieces (Sume docs and catalog, read 2026-10-09)
ApproachJobsCost
One split with 20 ranges1$0.01
20 splits with one range each20$0.20
Detach a video, then one split2$0.02

When to use mp3

Use mp3 when the pieces are final deliverables and size matters, such as voice lines for a game or app. Each mp3 edge adds priming padding, so do not join mp3 pieces again, and do not feed them to lip-sync. Keep the default wav for anything that will be joined, trimmed, or used as a timeline spine.

Offsets in segments[] (index, start, duration_seconds) let you line the pieces up with video slots. For ranges that overlap, such as a line with a one-second lead-in, each range still gets its own file.

Limits and errors

A split needs url and ranges (audio_split_requires_url, audio_split_requires_ranges). Concat parts and split ranges are capped at 20 each, so a 45-line script needs three split jobs, 3 x $0.01 = $0.03. Off-host URLs fail at admit with unsupported_media_source, and dead links with source_not_found. Provider or ffmpeg keys such as filtergraph or codec are rejected with a 400.

Concat is the reverse

The same endpoint with operation: "concat" joins up to 20 parts into one gapless file. The join is in the sample domain, with no re-synthesis and no silence at the seams, and the result's segments[] give offsets you use to re-base timeline video[].start. All parts must share a channel layout (audio_parts_channel_mismatch). If you only need the join inside a single render, use audio.parts[] on the timeline render instead; this job is for a reusable file.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume