Which Sume video models accept an audio reference (audio_url)?
Seedance 2.x, Wan 3.0, MiniMax H3 and H3 Max take audio and video references on Sume. Gemini Omni Flash 1.1 and Recast do not. Rates and limits.

On the Sume video API, a model lists the reference types it accepts in supported_input_references. The ones that matter for sound are audio_url and video_url. The video generation docs say the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references, while Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. That is the whole answer; the rest of this page is the table you need to act on it.
The reason to care is timing. Magic Hour's tracker (read 2026-10-07) describes Seedance 2.5 as making roughly 30-second clips in one pass with image, clip and text references, and lists LTX-2.5 with an audio-to-video endpoint. On Sume you reach the audio-reference job through the models below, not through LTX-2.5, which this page does not claim is in the Sume catalog.
Audio reference support, read 2026-10-07
| Model id | audio_url reference | Duration | Sume rate (list x 1.25) |
|---|---|---|---|
| seedance-2.5 | Yes | 4 to 30 s | By video tokens; see GET /v1/videos/models |
| wan-3.0 | Yes | 2 to 30 s | $0.0625/s 480p, $0.125/s 720p, $0.25/s 1080p |
| minimax-h3 | Yes | 5 to 15 s | $0.0625/s 480p, $0.075/s 768p |
| minimax-h3-max | Yes | 5 to 15 s | $0.0625/s 480p, $0.10/s 768p, $0.20/s 1080p |
| gemini-omni-flash-1.1 | No (image and video only) | 3 to 10 s | $0.0375/s 360p, $0.125/s 720p, $0.1875/s 1080p, $0.375/s 4K |
| higgsfield-genjutsu | No | 4 to 30 s | See catalog |
| h3-max-recast | No | 5 to 30 s | See catalog |
Ask the catalog, not this page
Limits differ for each model, so check the live list before you build around one. GET /v1/videos/models returns each model's supported_input_references, its generate_audio flag, durations and resolutions; the docs tell you to read it before you submit.
A face that has to speak a specific script is a different job. Even among the models that accept an audio reference, Sume routes on-camera speech to Fabric or H3 Max lip sync with a recorded voice, so the audio you supply drives the mouth exactly.
curl https://api.sume.com/v1/videos/models \
-H "Authorization: Bearer $SUME_API_KEY"How to choose
If your audio is a music bed or a sound design you already own and the video should move with it, take one of the four audio-aware models. If you only need the clip to carry its own sound, Omni Flash 1.1 makes native synced audio at $0.125 per second for 720p; just do not send it an audio_url. If you need a mouth on a specific voice, use the lip-sync routes in the spoken-line comparison.
Sources
Related posts
More in Models
- Which Sume video models take an audio_url reference, which do not
Sume video rows list supported_input_references per model. The docs show seedance-2 with audio_url, H3 Max with audio, and Gemini Omni Flash 1.1 without it.
- YouTube Shorts max out at 1080p, so don't generate at 4K
YouTube's Shorts help page caps uploads at 1080p. Sume's Gemini Omni Flash 1.1 offers 360p to 4K in 9:16, so request resolution 1080p and skip the higher tier.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
Written by Sume