Pull the audio out of a MiniMax H3 clip with Sume audio-detach
A minimax-h3 clip has native stereo audio. audio-detach returns it as a wav or mp3 artifact without touching the video. Fields, defaults and caps.

To get the soundtrack of a MiniMax H3 clip as a separate file, call POST /v1/audio-detach with the clip's media.sume.com URL. The default output is a sample-exact wav (pcm_s16le); you can ask for mp3 instead. The video is untouched. H3 and H3 Max generate native stereo audio with no toggle, so there is a track to detach.
Fields come from Sume's Audio detach docs, and H3's native audio from the Video Router docs, read 2026-10-02.
What does audio-detach need and return?
Input is one workspace media.sume.com video; there is no open-internet fetch, so a generated clip's result URL works and any other file must be imported first with POST /v1/media-imports. Idempotency-Key is required. The result is kind: audio_detach with an audio_url, duration_seconds, format, channels and sample_rate.
The result is a new artifact on a durable media.sume.com URL.
| Field | Notes |
|---|---|
video_url | Required; workspace media.sume.com artifact or asset |
format | Default wav; mp3 available |
range | Optional range to extract |
channels | Optional; sample_rate is inherited from the source when omitted |
Idempotency-Key | Required header |
Why would I detach the H3 audio?
Three common reasons: to reuse a generated sound bed under other footage, to send the speech to transcription, or to feed it into Timeline as an audio spine. Timeline's audio.url wants a wav, which is why the default is wav.
If you only need words on screen, skip the detach and use captions on the H3 clip directly.
Are there limits I should plan for?
The docs state the job is async by default and that mode: "sync" waits at most 30 seconds. H3 clips are 5 to 15 seconds, so a detach is small. Longer sources have their own caps; see the 900-second output cap.
Detach copies the audio as generated. It does not clean it up, and it does not change the picture or its timing, so keep the clip if you plan to rejoin them.
One check before you build on the track: play the generated audio first. Native audio is whatever the model made for that clip, and detaching it does not fix timing or loudness. Decide whether you need it before you spend on a second job.
Sources
Related posts
More in Media tools
- MiniMax H3 reference formats: HEIC, MOV, WAV, MP3 limits
MiniMax H3 accepts HEIC/HEIF images, MP4/MOV video and WAV/MP3 audio as references. A pre-flight checklist before you submit minimax-h3 on Sume.
- Keep the sound of AI clips in a Sume timeline: detach to a spine
A Timeline 1.0 render takes one audio spine. To keep audio generated with your clips, detach each clip's track, join the parts, and use that as the spine.
- Pika SFX: 1-20 second clips, negative prompts and seed vs Sume
Pika SFX makes 1 to 20 second effects with negative prompts, seed and guidance. Sume lists no sound-effects model, so here is what to use for each need.
- Fix a product ad's white balance with colortemperature in Sume
Warm or cool a product clip with ffmpeg colortemperature in Sume video filter: 1000 to 40000 K, default 6500, plus mix and pl. Dry-run it for free.
Written by Sume