Which AI video models take a reference audio file on Sume?
Seedance 2.x, Wan 3.0, MiniMax H3 and H3 Max accept a reference audio clip on Sume; Gemini Omni Flash, Kling, Grok, Recast and Genjutsu do not.

A reference audio file is different from the audio a model generates. Native audio, such as Gemini Omni Flash's always-on soundtrack, is something the model makes up. A reference audio is a clip you supply, for the model to use as a guide for voice, rhythm or sound. Only some Sume rows accept one, and the rest return a 400 if you send it.
The split comes from each row's capability flags in the Video Router catalog, which GET /v1/videos/models projects as supported_input_references; a row that lists audio_url there accepts a reference audio.
Which rows accept a reference audio?
Seedance 2.5, the Seedance 2.0 family, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept one. Gemini Omni Flash 1.1, Kling Video v3 Pro, Grok Imagine Video 1.5, H3 Max Recast and Higgsfield Genjutsu do not.
| Row | Accepts reference audio |
|---|---|
| Seedance 2.5 and Seedance 2.0 family | Yes |
wan-3.0 | Yes, up to 5 clips, 15 s total |
minimax-h3 | Yes |
minimax-h3-max | Yes |
gemini-omni-flash-1.1 | No |
kling-3 | No |
grok-imagine-video-1.5 | No |
h3-max-recast | No |
higgsfield-genjutsu | No |
What are the limits?
Wan 3.0 caps reference audio at five clips and 15 seconds in total. The other accepting rows do not publish a count in the constraints Sume exposes, so keep to one short clip unless you have tested more. Recast and Genjutsu keep the soundtrack of the source video as is, so they need no audio input.
Gemini Omni Flash 1.1 and Kling also differ on the generated-audio switch: Omni always produces audio and has no toggle, so generate_audio: false is rejected there and on sume/auto, while Kling bills audio on and off at different rates.
How do I send one?
On /v1/videos, add an input_references entry of type audio_url alongside any image or video references. On the legacy Video Router wire the audio goes in the matching reference field. If a row does not list audio_url, remove the entry rather than hoping it is ignored; Sume validates against the catalog and does not silently drop fields.
See Seedance 2.5 input references for the JSON shapes.
Sources
Related posts
More in Models
- Which Sume image models accept quality? Only five do
Only five Sume image catalog rows list a quality field: GPT Image 2, 2.5 and Sunburst, Ideogram V3 and 4.5. The rest return 400 unsupported_parameter.
- Which Sume image models can't edit? Five text-to-image-only rows
Five Sume image rows take text prompts only and reject image_urls: Soul, Imagen 4 Fast and Ultra, Recraft V4 and Qwen Image Max. Prices and what to use instead.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
- Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.
Written by Sume