Pull a voice track from AI video, then join takes gaplessly
Extract a wav or mp3 from a Sume-hosted video with POST /v1/audio-detach, then join up to 20 audio parts with Timeline audio. Fields, caps, $0.01 rate.

Send one Sume-hosted video to POST /v1/audio-detach and you get a new audio artifact back. The default is a sample-exact pcm_s16le wav, and the video is untouched. To join several audio takes into one file with no gap, send up to 20 ordered parts to POST /v1/timeline-1.0/audio with operation: "concat".
Extract the track
Audio detach takes one video_url that must be this workspace's media.sume.com artifact or asset. It does not fetch from the open internet, so import first with POST /v1/media-imports. An Idempotency-Key header is required, and the default mode is async.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audio-detach-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4"}'Read the result
There is no GET /v1/audio-detach/:id. Poll GET /v1/jobs/:id/status, then read GET /v1/jobs/:id/result when result_ready is true. The result has audio_url, duration_seconds, format, channels, sample_rate and source_duration_seconds.
| Field | Values | Note |
|---|---|---|
format | wav (default) or mp3 | wav is pcm_s16le; mp3 is 128 kbps |
range | { start, end? } in seconds | Omit for the whole track |
channels | source (default) or mono | |
sample_rate | 16000, 44100 or 48000 | Omit to inherit; 16000 mono is the speech-to-text shape |
Caps and refusals
The source can be at most 1800 seconds and the output at most 900 seconds, so a longer track needs a range. A video with no audio fails with detach_source_has_no_audio; check probe.has_audio with Video inspect first. Do not send ffmpeg fields such as af or codec: the server compiles ffmpeg itself and answers ffmpeg_fields_rejected.
Join takes without a gap
Timeline audio concat takes parts[] of 1 to 20 items, each { url, source_in?, duration? }, and joins them in the sample domain with no re-synthesis and no silence at the seams. Keep wav, which is the default, when the file will be joined again. mp3 re-adds priming padding at every edge.
The result gives one audio_url plus segments[] with index, start and duration_seconds. Those are the offsets to use when you place video slots in a Timeline render. Parts must share one channel layout or you get audio_parts_channel_mismatch.
Cost and the many-ranges case
Both jobs are listed at $0.01 each, with no provider inference, only worker ffmpeg. Confirm the live rate in GET /v1/catalog. If you need many ranges from one talking video, detach once and then use Timeline audio with operation: "split" rather than detaching repeatedly.
Sources
Related posts
More in Media tools
- Swapped the voice? Re-time the visuals from word timestamps
A new voice speaks at a new pace, so cuts set for the old one drift. Take each scene's start from the new take's word timestamps and re-plan the timeline.
- Reference ingest coverage: how many frames OCR read
The reference-ingest coverage block reports frames decoded, frames OCR read and the OCR rate. Use it to decide when on-screen text needs a second look.
- Reference ingest purpose: qa or remix decides who transcribes
Reference ingest purpose defaults to reference_remix, which transcribes speech. brief_format, face_swap and qa do not. An explicit allow_billed_stt wins.
- Reference ingest shot record: cut types, palette, luma, motion
Each reference-ingest shot carries cut_out type (hard, gradual, end), palette, luma, contrast and a motion class. Turn them into shot length and pacing numbers.
Written by Sume