Pull the audio out of a MiniMax H3 clip with Sume audio-detach

A minimax-h3 clip has native stereo audio. audio-detach returns it as a wav or mp3 artifact without touching the video. Fields, defaults and caps.

4 min readSume
All posts

To get the soundtrack of a MiniMax H3 clip as a separate file, call POST /v1/audio-detach with the clip's media.sume.com URL. The default output is a sample-exact wav (pcm_s16le); you can ask for mp3 instead. The video is untouched. H3 and H3 Max generate native stereo audio with no toggle, so there is a track to detach.

Fields come from Sume's Audio detach docs, and H3's native audio from the Video Router docs, read 2026-10-02.

What does audio-detach need and return?

Input is one workspace media.sume.com video; there is no open-internet fetch, so a generated clip's result URL works and any other file must be imported first with POST /v1/media-imports. Idempotency-Key is required. The result is kind: audio_detach with an audio_url, duration_seconds, format, channels and sample_rate.

The result is a new artifact on a durable media.sume.com URL.

Audio detach request fields from Sume's Audio detach docs, read 2026-10-02.
FieldNotes
video_urlRequired; workspace media.sume.com artifact or asset
formatDefault wav; mp3 available
rangeOptional range to extract
channelsOptional; sample_rate is inherited from the source when omitted
Idempotency-KeyRequired header

Why would I detach the H3 audio?

Three common reasons: to reuse a generated sound bed under other footage, to send the speech to transcription, or to feed it into Timeline as an audio spine. Timeline's audio.url wants a wav, which is why the default is wav.

If you only need words on screen, skip the detach and use captions on the H3 clip directly.

Are there limits I should plan for?

The docs state the job is async by default and that mode: "sync" waits at most 30 seconds. H3 clips are 5 to 15 seconds, so a detach is small. Longer sources have their own caps; see the 900-second output cap.

Detach copies the audio as generated. It does not clean it up, and it does not change the picture or its timing, so keep the clip if you plan to rejoin them.

One check before you build on the track: play the generated audio first. Native audio is whatever the model made for that clip, and detaching it does not fix timing or loudness. Decide whether you need it before you spend on a second job.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume