Which Sume video models accept an audio reference (audio_url)?

Seedance 2.x, Wan 3.0, MiniMax H3 and H3 Max take audio and video references on Sume. Gemini Omni Flash 1.1 and Recast do not. Rates and limits.

5 min readSume
All posts

On the Sume video API, a model lists the reference types it accepts in supported_input_references. The ones that matter for sound are audio_url and video_url. The video generation docs say the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references, while Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. That is the whole answer; the rest of this page is the table you need to act on it.

The reason to care is timing. Magic Hour's tracker (read 2026-10-07) describes Seedance 2.5 as making roughly 30-second clips in one pass with image, clip and text references, and lists LTX-2.5 with an audio-to-video endpoint. On Sume you reach the audio-reference job through the models below, not through LTX-2.5, which this page does not claim is in the Sume catalog.

Audio reference support, read 2026-10-07

Video models and audio references on Sume, read 2026-10-07
Model idaudio_url referenceDurationSume rate (list x 1.25)
seedance-2.5Yes4 to 30 sBy video tokens; see GET /v1/videos/models
wan-3.0Yes2 to 30 s$0.0625/s 480p, $0.125/s 720p, $0.25/s 1080p
minimax-h3Yes5 to 15 s$0.0625/s 480p, $0.075/s 768p
minimax-h3-maxYes5 to 15 s$0.0625/s 480p, $0.10/s 768p, $0.20/s 1080p
gemini-omni-flash-1.1No (image and video only)3 to 10 s$0.0375/s 360p, $0.125/s 720p, $0.1875/s 1080p, $0.375/s 4K
higgsfield-genjutsuNo4 to 30 sSee catalog
h3-max-recastNo5 to 30 sSee catalog

Ask the catalog, not this page

Limits differ for each model, so check the live list before you build around one. GET /v1/videos/models returns each model's supported_input_references, its generate_audio flag, durations and resolutions; the docs tell you to read it before you submit.

A face that has to speak a specific script is a different job. Even among the models that accept an audio reference, Sume routes on-camera speech to Fabric or H3 Max lip sync with a recorded voice, so the audio you supply drives the mouth exactly.

curl https://api.sume.com/v1/videos/models \
  -H "Authorization: Bearer $SUME_API_KEY"

How to choose

If your audio is a music bed or a sound design you already own and the video should move with it, take one of the four audio-aware models. If you only need the clip to carry its own sound, Omni Flash 1.1 makes native synced audio at $0.125 per second for 720p; just do not send it an audio_url. If you need a mouth on a specific voice, use the lip-sync routes in the spoken-line comparison.

Sources

Related posts

More in Models

All Models posts

Written by Sume