Audio-only cut of a 5 minute episode: mp3 for $0.01
Turn a finished 5 minute video into a 128 kbps mp3 for $0.01 with audio-detach. Source up to 1800 s, output up to 900 s, and the silent-source error.

Audio detach turns a finished episode into an mp3 for a flat $0.01 per job, whether the video is 60 seconds or 5 minutes. The source may run to 1800 seconds and the output to 900 seconds. Use it for a listen-only version, a transcript source or a voice-note cut when you already have the video.
TikTok is reported by Metricool to favour videos of 3 to 5 minutes. An audio-only copy gives you a second asset from the same expensive render.
The request
Send the video_url and format. The default is wav (pcm_s16le); mp3 is 128 kbps. channels can be source or mono, and sample_rate can be 16000, 44100 or 48000. A range of start and optional end seconds trims the audio.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ep3-audio-mp3" \
-d '{"video_url":"https://media.sume.com/artifacts/artf_ep3/ep3.mp4","format":"mp3"}'Limits that matter
The result is a new artifact with audio_url, duration_seconds, format, channels and source_duration_seconds. It is worker ffmpeg, so no model inference.
| Item | Value |
|---|---|
| Price | $0.01 per job |
| Source length | Up to 1800 s |
| Output length | Up to 900 s |
| Formats | wav (default), mp3 at 128 kbps |
| Silent source | Fails with detach_source_has_no_audio |
Check for sound first
A clip with no audio track fails, and a failed job should not surprise you. Run video inspect with frames set to false to read probe.has_audio without making stills. Generated clips from a model with audio off, and silent timeline renders, are the usual cases.
A 5 minute episode at 900 seconds of output headroom is well inside the cap, so no splitting is needed. Longer sources use range to take pieces.
Where it fits
Pair it with captions and a transcript pass for search text, or hand the file to a podcast host. Sume does not publish to feeds or add chapters. It makes the file; where it goes is yours to decide.
Sources
Related posts
More in Media tools
- Batch trim clips from a spreadsheet of start and end times
Read start and end columns from a CSV and trim a long video into Shorts: one Video trim call per row at $0.02, with idempotency keys.
- Black bars on a YouTube Short: YouTube says no, so fill the frame
YouTube says uploads should never include letterbox or pillarbox bars, and Shorts take square or vertical files. Reframe 16:9 clips with Sume Timeline fit.
- Caption a Recast video: burn captions after the person swap
Run h3-max-recast first, then send its output to POST /v1/video-captions. The kept source audio gives the captions speech to transcribe.
- Caption micro-drama dialogue per shot: $0.20 per clip up to 60 s
Video captions cost $0.20 per job for clips up to 60 seconds. Caption each shot before assembly; a silent clip fails with caption_no_speech. Budget math.
Written by Sume