AI music video generator API: which models take your audio?
Seedance, Wan and MiniMax accept audio references on Sume. Gemini Omni and Kling 3 do not. Limits per model for a music-video workflow.

Which video models can use your own song as a reference? On Sume, the Seedance 2 family, Seedance 2.5, Wan 3.0 and the MiniMax H3 models accept audio references. Gemini Omni Flash 1.1, h3-max-recast and higgsfield-genjutsu do not, and the kling-3 row has no reference fields. Clip length is the next limit: one pass reaches 30 seconds on Seedance 2.5 and Wan 3.0.
| Model on Sume | Audio reference | Clip length | Audio limit |
|---|---|---|---|
| seedance-2.5 | Yes | 4-30 s | Not read |
| wan-3.0 | Yes | 2-30 s | Up to 5 clips, 15 s total |
| minimax-h3 / minimax-h3-max | Yes | 5-15 s | Up to 3 clips, 15 s total |
| gemini-omni-flash-1.1 | No | 3-10 s | Native audio only |
| kling-3 | No reference fields | 4-15 s | Audio on/off pricing |
Vendor context
The Seed announcement lists up to 10 audio clips per pass for Seedance 2.5. The MiniMax guide gives 3 audio clips of 2 to 15 seconds, 15 seconds total, WAV or MP3 up to 15 MB. The Kling 4.0 post describes voice references, but Kling 4.0 is not on Sume.
How a song fits
A three minute song does not fit any one clip, so cut the track into sections and generate a clip per section, then assemble them on a timeline with the full track as the audio spine. Use video-trim for video sources; for the audio, cut sections with your own tools before upload.
I did not test how closely any model follows rhythm, and the vendor pages I read do not promise beat sync, so judge each model on your track.
Cost sketch
For a 30 second section, Wan 3.0 lists $0.05, $0.10 and $0.20 per second at 480p, 720p and 1080p, which Sume bills at list times 1.25 rounded up to cents. Seedance 2.5 bills from a per-1,000-video-token rate, so confirm the price in pricing_skus before a batch.
If your plan needs the model itself to produce the soundtrack, Omni and Kling 3 generate audio natively but do not take a track as a reference, so they suit clips where the sound can be generated and replaced on the timeline.
Check the frame rate and format of any reference before upload. Wan 3.0 requires reference videos of at least 16 fps on Sume, and MiniMax takes WAV or MP3 audio up to 15 MB. A source that fails these limits will fail the request, so prepare files first.
Sources
Related posts
More in Use cases
- AI product demo avatar: live agent or recorded clip?
Tavus builds live demo agents that answer buyers; Sume renders a 4-60 second avatar clip with your product image. Where each fits, plus Sume per-second prices.
- An AI series render ledger: job id, key, plan minutes, warnings
TikTok is testing spam detection on AI-content accounts. Keep a CSV of every Sume render: job id, idempotency key, plan minutes, status, file and warnings.
- AI video generator for kids: what YouTube says and what to build
Over 200 groups asked YouTube to ban AI videos for kids. What YouTube replied, and how to use Sume for family or classroom clips without chasing views.
- Replace the voice in a video you shot: detach, transcribe, re-voice
Swap the speech in a finished video on Sume: detach the audio, transcribe it, fix the script, speak it with TTS and render the picture with the new track.
Written by Sume