Split one recording into many audio clips: detach once, then split
Twenty ranges from one talking-head video need one audio detach and one split job, not twenty cuts. Wav or mp3, overlap rules, and what each clip returns.

How do you cut one long recording into several audio clips through an API? Detach the audio track once, then call timeline audio with operation: "split" and up to 20 ranges. Each range comes back as its own durable file, so you do not repeat a media job per cut.
Streaming transcribers such as Microsoft's MAI-Transcribe-2-Streaming (read 2026-10-04) work on a live feed. Segmenting a recording into clips is the batch counterpart: prepare the pieces, then send each to transcription or captions.
Step 1 and step 2
Audio detach turns one media.sume.com video into a new audio artifact, wav by default (pcm_s16le, sample-exact). The video is untouched. Output is capped at 900 seconds, so a longer track needs a range.
Timeline audio then slices that file. Required are operation: "split", a top-level url and ranges[] (1 to 20), each { start, end? }; omit end to mean the rest of the file. Do not send parts (audio_split_takes_no_parts). Ranges may overlap, which suits overlapping pull quotes.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: split-quotes-001" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/spine.wav",
"ranges": [
{ "start": 12.4, "end": 31.0 },
{ "start": 95.0, "end": 118.5 },
{ "start": 240.0 }
]
}'
What you get and what it costs
The result is kind: timeline_audio with segments[], each carrying its own audio_url.
| Item | Detail |
|---|---|
| Detach price | $0.01 per job |
| Split price | $0.01 flat per job |
| Ranges per split | 1 to 20 |
| Default format | wav, sample-exact |
| mp3 option | Smaller; re-adds priming padding at every edge |
Choose the format by what comes next
Keep wav when a clip will be joined again or will drive lip-sync. Use mp3 only for delivery files, because it re-adds padding at each edge. If you plan to transcribe each clip, a 16 kHz mono wav from the detach step is the STT shape.
More than 20 ranges means two split jobs against the same detached file. Produced audio is capped at 1800 seconds per job.
Sources
Related posts
More in Media tools
- Stability AI's Series B and label backers: what it means for audio
Stability released Stable Audio 3.0 on 5/20/26 and raised a Series B on 8/25/26 with EA, Sony, UMG and WMG. A dated timeline and what to verify for video work.
- Stabilize shaky video by API: Sume has no deshake, so do this
Sume's video filter refuses deshake and vidstab as unknown filters. Prove it with the free check, stabilize upstream, then use Sume for trim, crop and captions.
- Temu detail video up to 180 s: stitch 30 s clips with Timeline
Temu detail video allows up to 180 s and 300 MB at 720p or higher. Sume clips top out at 30 s, so six clips stitched with Timeline cover the full length.
- TikTok trending API: browse without a query is flag-gated
Sume trending search needs a query on production. Browse with no query, limit up to 100 and relevance floors are dev-only until a flag is set. The differences.
Written by Sume