Podcast guest clips from a video call: detach once, then split
Five audio clips from one recorded call cost $0.02 on Sume: one audio-detach ($0.01) plus one timeline-audio split ($0.01), not $0.05 for five detaches.

To pull five quotable clips from a recorded video call, run audio detach once on the whole recording, then one timeline-audio split with five ranges. That is $0.01 plus $0.01, or $0.02, against 5 x $0.01 = $0.05 if you detach each range separately. Both jobs are flat-rate, ffmpeg-only, with no provider inference.
Sources: Audio detach and Timeline audio, read 2026-10-04. The audio-detach page itself recommends this order for many ranges.
The four calls
The path is import, check, detach, split. A whole-track detach is limited to 900 seconds of output, so a longer call needs a range on the detach job.
| Step | Route | Price |
|---|---|---|
| Import the call recording | POST /v1/media-imports | Required before any Sume job reads it |
| Check there is audio | POST /v1/video-inspect with frames false | Billed by Modal compute |
| Detach the track | POST /v1/audio-detach | $0.01 per job |
| Split into 5 clips | POST /v1/timeline-1.0/audio, operation split | $0.01 per job |
Detach once
Detach with the default wav output so the split is sample-exact. The result carries audio_url, duration_seconds and the probed source duration. If the file has no track you get a stable refusal rather than a silent file, which is why the probe step is worth the cost.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: guest-call-detach-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/guest-call.mp4",
"format": "wav"
}'Split into the five quotes
The split takes up to 20 ranges, each { start, end } in seconds; end omitted means the rest of the file, and ranges may overlap. The result is a segments[] list, each with its own audio_url.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: guest-call-split-001" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/guest-call-audio.wav",
"ranges": [
{ "start": 62.0, "end": 88.5 },
{ "start": 240.0, "end": 271.2 },
{ "start": 455.8, "end": 480.0 },
{ "start": 702.4, "end": 731.0 },
{ "start": 811.0, "end": 840.6 }
]
}'Where the ranges come from
Find the ranges by transcribing the call first: video-inspect with transcribe: true and segmentation.mode: "sentence" returns gapless sentence segments[] with timings. The transcript is $0.01 per audio minute, with a duration_seconds hint of at most 600. Pick sentence boundaries as range edges so no quote starts mid-word.
Sources
Related posts
More in Use cases
- Animate podcast cover art into a season-launch clip on Sume
Turn your podcast cover into a 6-second motion clip for the season launch: Wan 3.0 first frame at 720p is $0.75, then add audio with one Timeline render.
- Podcast teaser clip: cover art plus audio reference on Seedance 2.5
Make a 15 to 30 s teaser by sending cover art and an episode excerpt as reference_audio_urls to Seedance 2.5 on Sume. Honest lip-sync limits and cost.
- Price increase announcement video: a 45-second avatar, four scenes
Announce a price change with a 45-second Sume avatar video in four scenes: hook, reason, what changes, next step. Cost by tier, a working video_inputs body.
- Add print bleed to an AI image: mirror padding in Pillow
An AI image has no bleed, so a trimmed print can show a white edge. Extend the edges by mirroring with NumPy and Pillow, to the bleed your printer specifies.
Written by Sume