Split a 14-minute episode into 20 clips with overlaps: 2 cents
Detach the audio once ($0.01), then one timeline audio split with 20 overlapping ranges ($0.01). A 860-second track fits the 900-second detach cap.

Detach the episode's audio once for $0.01, then send a single timeline audio split with up to 20 ranges for another $0.01: 20 clips from a 14-minute-20-second (860-second) episode cost $0.02 in total. Ranges may overlap, which is how you give each clip a 1-second handle on both sides.
Everything here is from the Sume audio detach and timeline audio docs and the API catalog (read 2026-10-09). The tools only run ffmpeg on Sume's workers; neither uses provider inference.
Why 860 seconds works
Audio detach caps the source at 1,800 seconds and the output at 900 seconds. A whole track longer than 900 seconds needs a range, so 860 seconds fits as a whole-track wav. Twenty clips from it average 43 seconds each (860 / 20). With a 1-second handle on each side, the ranges overlap their neighbors by 1 second.
Range i could be { "start": 43*i - 1, "end": 43*(i+1) + 1 }, clamped at 0 for the first clip and left without an end on the last. The docs say that omitting end goes to the end of the file, and overlapping ranges are allowed.
| Step | Endpoint | Limit that matters | Cost |
|---|---|---|---|
| Detach audio to wav | POST /v1/audio-detach | Output at most 900 s | $0.01 |
| Split into clips | POST /v1/timeline-1.0/audio, operation split | 1 to 20 ranges, may overlap | $0.01 |
| Total | $0.02 |
Wav, not mp3
Keep wav (the default, sample-exact) if the clips will be trimmed or joined again. The docs note that mp3 adds priming padding at every edge, and a clip with a 1-second handle is meant to be trimmed again later.
A twenty-first clip needs a second split job, $0.01 more. Plan the 20 ranges to cover the sections you really need, and use the last range open-ended if you want the tail of the episode.
Reading the result
The split result is kind: timeline_audio with segments[], and each segment has its own audio_url. Because the job is asynchronous by default, poll GET /v1/jobs/:id/status or pass mode: "sync" to wait up to 30 seconds. Send an Idempotency-Key on both calls so a retry does not bill twice.
What to check
Before you publish the clips, compare the reported duration_seconds of each segment with the range you asked for. The handle overlap means neighboring clips share a second of audio, so do not re-join them by simple concatenation or you will repeat that second.
If the source is a video and has no audio track, the detach fails with detach_source_has_no_audio and you are not charged for the split you never submitted. A frames: false inspect tells you probe.has_audio first.
Both calls need an Idempotency-Key, and the tools are also exposed on hosted MCP as audio_detach and timeline_audio, which write with an idempotency_key. The flow there is the tool, then jobs_wait, then jobs_result.
Sources
Related posts
More in Use cases
- Sponsored Snaps headline: 24 to 28 characters, check a batch
Snapchat's Sponsored Snaps page recommends a 24-28 character headline on a 9:16 asset. Check a batch of headlines in Python before you queue the Sume clips.
- Square 1:1 AI video for YouTube Shorts: Seedance 2.0 15 s from $2.64
YouTube counts square or vertical uploads up to 3 minutes as Shorts. A 15-second 1:1 Seedance 2.0 clip on Sume is $2.64 at 480p, $5.67 at 720p, $12.76 at 1080p.
- Square product photo to a 4:1 banner: reference plus aspect_ratio
Turn a 1:1 photo into a 4:1 banner with Nano Banana 2.1 on Sume: input_references plus an explicit aspect_ratio of 4:1, not auto. $0.10 at 1K. Body and gotchas.
- Support answer video with an avatar: 25 seconds with a product shot
A 25-second support answer video costs $4.60 on Standard, $6.125 on Plus and $13.75 on Max. A product image adds a small per-second premium.
Written by Sume