Split one recording into many audio clips: detach once, then split

Twenty ranges from one talking-head video need one audio detach and one split job, not twenty cuts. Wav or mp3, overlap rules, and what each clip returns.

5 min readSume
All posts

How do you cut one long recording into several audio clips through an API? Detach the audio track once, then call timeline audio with operation: "split" and up to 20 ranges. Each range comes back as its own durable file, so you do not repeat a media job per cut.

Streaming transcribers such as Microsoft's MAI-Transcribe-2-Streaming (read 2026-10-04) work on a live feed. Segmenting a recording into clips is the batch counterpart: prepare the pieces, then send each to transcription or captions.

Step 1 and step 2

Audio detach turns one media.sume.com video into a new audio artifact, wav by default (pcm_s16le, sample-exact). The video is untouched. Output is capped at 900 seconds, so a longer track needs a range.

Timeline audio then slices that file. Required are operation: "split", a top-level url and ranges[] (1 to 20), each { start, end? }; omit end to mean the rest of the file. Do not send parts (audio_split_takes_no_parts). Ranges may overlap, which suits overlapping pull quotes.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: split-quotes-001" \
  -d '{
    "operation": "split",
    "url": "https://media.sume.com/artifacts/artf_demo/spine.wav",
    "ranges": [
      { "start": 12.4, "end": 31.0 },
      { "start": 95.0, "end": 118.5 },
      { "start": 240.0 }
    ]
  }'

What you get and what it costs

The result is kind: timeline_audio with segments[], each carrying its own audio_url.

Detach plus split, read 2026-10-04
ItemDetail
Detach price$0.01 per job
Split price$0.01 flat per job
Ranges per split1 to 20
Default formatwav, sample-exact
mp3 optionSmaller; re-adds priming padding at every edge

Choose the format by what comes next

Keep wav when a clip will be joined again or will drive lip-sync. Use mp3 only for delivery files, because it re-adds padding at each edge. If you plan to transcribe each clip, a 16 kHz mono wav from the detach step is the STT shape.

More than 20 ranges means two split jobs against the same detached file. Produced audio is capped at 1800 seconds per job.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume