Reference audio for AI video: which models accept a voice clip

Seedance 2.5, Wan 3.0 and MiniMax H3 take reference audio on Sume; Kling 3, Gemini Omni Flash and the swap rows do not. Limits and the one-reference rule.

4 min readSume
All posts

Reference audio works on Seedance 2.5 and the Seedance 2.0 rows, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Kling 3 has no reference inputs at all, Gemini Omni Flash 1.1 takes image and video references but not audio, and Genjutsu and H3 Max Recast refuse audio references. Where it is accepted, audio cannot be the only reference: send at least one reference image or video with it.

Limits by model

Counts below are the catalog constraints for the rows that publish them. For Seedance, see the linked post on how reference types mix, since the cap there is a total rather than a per-type number.

Reference audio support, read 2026-10-03. Source: Sume video catalog on origin/main, read 2026-10-03.
Model idReference audioLimit
wan-3.0yesup to 5 clips, 15 s total
minimax-h3, minimax-h3-maxyesup to 3 clips, 2 to 15 s each, 15 s combined; images, videos and audios total 12
seedance-2.5yesneeds an image or video reference alongside
kling-3noreference_*_urls rejected
gemini-omni-flash-1.1nono reference_audio_urls
genjutsu, h3-max-recastnoaudio references rejected

Using it

Reference audio is for matching a voice, a beat or a mood, not for pasting a finished soundtrack under the clip. For a finished track, generate the video and mix the music on a timeline afterwards.

Audio URLs must be public HTTPS. MiniMax and Wan enforce duration limits on each clip, so trim an audio file before submitting rather than relying on the model to cut it.

A quick check before you submit

If your request has only reference_audio_urls, it will be refused. Add a first-frame image or a reference image, then retry.

Sources

Related posts

More in Models

All Models posts

Written by Sume