Make AI Video Follow Your Audio: Which Models Take Audio References

Seedance, Wan 3.0 and MiniMax H3 accept an audio reference on Sume; Omni does not. Request shape, Wan limits and what an audio reference is not.

5 min readSume
All posts

Some video models can take a sound file as an input and let it shape the clip, which is different from lip sync. Sume's docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, while Gemini Omni Flash 1.1, Genjutsu and H3 Max Recast accept video references but not audio.

Limits that are published

Wan 3.0 on Sume takes up to 5 audio references with a combined length of at most 15 seconds. ByteDance says Seedance 2.5 takes up to 10 audio clips and generates audio and video jointly (read 2026-10-03). Sume's docs do not publish an audio count for the Seedance rows or for H3, so treat the vendor's number as an upper bound. Kling 3 on Sume accepts no references at all.

Audio references, read 2026-10-03
ModelAudio referencePublished limit
seedance-2.5 and 2.0 familyYesNot published by Sume; vendor says 10 for 2.5
wan-3.0YesUp to 5, 15 s total
minimax-h3 and minimax-h3-maxYesNot published
gemini-omni-flash-1.1NoNot accepted
kling-3NoNo references

The request

On POST /v1/videos, an audio reference is an entry in input_references with type: audio_url. This example uses Wan with one music bed.

{
  "model": "wan-3.0",
  "prompt": "Product shots cut to the rhythm of the music, warm light",
  "resolution": "720p",
  "duration": 10,
  "input_references": [
    {
      "type": "audio_url",
      "audio_url": { "url": "https://example.com/music-bed.mp3" }
    }
  ]
}

What it does not do

An audio reference guides the model; it is not a promise of beat-accurate cuts or word-accurate speech. Sume's docs say video models do not lip-sync to generated TTS or a later voice-over, and that a talking face belongs on the still-plus-audio routes. For exact beat timing, generate clips and set cut points in Timeline instead.

The only way to know how well a given track steers a given model is to test a 4-second clip. At $0.125 per second at 720p on Wan, that test costs about $0.50.

Sources

Related posts

More in Models

All Models posts

Written by Sume