Pull a voice track from AI video, then join takes gaplessly

Extract a wav or mp3 from a Sume-hosted video with POST /v1/audio-detach, then join up to 20 audio parts with Timeline audio. Fields, caps, $0.01 rate.

4 min readSume
All posts

Send one Sume-hosted video to POST /v1/audio-detach and you get a new audio artifact back. The default is a sample-exact pcm_s16le wav, and the video is untouched. To join several audio takes into one file with no gap, send up to 20 ordered parts to POST /v1/timeline-1.0/audio with operation: "concat".

Extract the track

Audio detach takes one video_url that must be this workspace's media.sume.com artifact or asset. It does not fetch from the open internet, so import first with POST /v1/media-imports. An Idempotency-Key header is required, and the default mode is async.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audio-detach-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4"}'

Read the result

There is no GET /v1/audio-detach/:id. Poll GET /v1/jobs/:id/status, then read GET /v1/jobs/:id/result when result_ready is true. The result has audio_url, duration_seconds, format, channels, sample_rate and source_duration_seconds.

Audio detach options (Sume docs) (read 2026-10-03)
FieldValuesNote
formatwav (default) or mp3wav is pcm_s16le; mp3 is 128 kbps
range{ start, end? } in secondsOmit for the whole track
channelssource (default) or mono
sample_rate16000, 44100 or 48000Omit to inherit; 16000 mono is the speech-to-text shape

Caps and refusals

The source can be at most 1800 seconds and the output at most 900 seconds, so a longer track needs a range. A video with no audio fails with detach_source_has_no_audio; check probe.has_audio with Video inspect first. Do not send ffmpeg fields such as af or codec: the server compiles ffmpeg itself and answers ffmpeg_fields_rejected.

Join takes without a gap

Timeline audio concat takes parts[] of 1 to 20 items, each { url, source_in?, duration? }, and joins them in the sample domain with no re-synthesis and no silence at the seams. Keep wav, which is the default, when the file will be joined again. mp3 re-adds priming padding at every edge.

The result gives one audio_url plus segments[] with index, start and duration_seconds. Those are the offsets to use when you place video slots in a Timeline render. Parts must share one channel layout or you get audio_parts_channel_mismatch.

Cost and the many-ranges case

Both jobs are listed at $0.01 each, with no provider inference, only worker ffmpeg. Confirm the live rate in GET /v1/catalog. If you need many ranges from one talking video, detach once and then use Timeline audio with operation: "split" rather than detaching repeatedly.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume