Make AI Video Follow Your Audio: Which Models Take Audio References
Seedance, Wan 3.0 and MiniMax H3 accept an audio reference on Sume; Omni does not. Request shape, Wan limits and what an audio reference is not.

Some video models can take a sound file as an input and let it shape the clip, which is different from lip sync. Sume's docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, while Gemini Omni Flash 1.1, Genjutsu and H3 Max Recast accept video references but not audio.
Limits that are published
Wan 3.0 on Sume takes up to 5 audio references with a combined length of at most 15 seconds. ByteDance says Seedance 2.5 takes up to 10 audio clips and generates audio and video jointly (read 2026-10-03). Sume's docs do not publish an audio count for the Seedance rows or for H3, so treat the vendor's number as an upper bound. Kling 3 on Sume accepts no references at all.
| Model | Audio reference | Published limit |
|---|---|---|
| seedance-2.5 and 2.0 family | Yes | Not published by Sume; vendor says 10 for 2.5 |
| wan-3.0 | Yes | Up to 5, 15 s total |
| minimax-h3 and minimax-h3-max | Yes | Not published |
| gemini-omni-flash-1.1 | No | Not accepted |
| kling-3 | No | No references |
The request
On POST /v1/videos, an audio reference is an entry in input_references with type: audio_url. This example uses Wan with one music bed.
{
"model": "wan-3.0",
"prompt": "Product shots cut to the rhythm of the music, warm light",
"resolution": "720p",
"duration": 10,
"input_references": [
{
"type": "audio_url",
"audio_url": { "url": "https://example.com/music-bed.mp3" }
}
]
}What it does not do
An audio reference guides the model; it is not a promise of beat-accurate cuts or word-accurate speech. Sume's docs say video models do not lip-sync to generated TTS or a later voice-over, and that a talking face belongs on the still-plus-audio routes. For exact beat timing, generate clips and set cut points in Timeline instead.
The only way to know how well a given track steers a given model is to test a 4-second clip. At $0.125 per second at 720p on Wan, that test costs about $0.50.
Sources
Related posts
More in Models
- Mercury Voice is enterprise-only: pricing and what to ask sales
Mercury Voice is GA for enterprise customers only. List $0.40/$1.50 per million tokens, launch price $0.20/$0.75. Questions to ask before you commit.
- Mercury Voice p95 750 ms: one turn in twenty is slower
Inception reports a 320 ms median and 750 ms p95 for Mercury Voice. At p95, about one turn in 20 is slower. What that means over a 10-turn call.
- MiniMax H3 2K and 4K upscale on Sume: H3 accepts them, H3 Max does not
minimax-h3 will price a 2K or 4K request even though its resolution list shows only 480p and 768p. minimax-h3-max rejects both. Costs for 5 to 15 seconds.
- MiniMax H3 or H3 Max on Sume: which id for a 5 to 15 second clip?
Choose minimax-h3 for a 480p or 768p draft at $0.075 a second, minimax-h3-max when you need 1080p. Same 5-15 s range, same references; Max costs more at 768p.
Written by Sume