Reference audio for AI video: which Sume models take audio_url
Seedance 2.x, Wan 3.0 and MiniMax H3 honor audio references on Sume; Gemini Omni, Genjutsu and Recast take video but not audio. Rules and a check.

On Sume, audio references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. Check supported_input_references for audio_url before you send one, because a model that does not list it rejects the field.
The question comes up with every launch that mentions audio. Kling's announcement for Kling 4.0 (read 2026-10-02) describes stereo audio with improved lip-sync, but that is audio the model produces, not audio you supply. Sume's docs treat supplied audio as a reference input, and that support is per model.
Which models honor an audio reference?
The table lists what Sume's video generation and Video Router docs state. Your live catalog is the final word.
| Model | Video reference | Audio reference |
|---|---|---|
| Seedance 2.x | Yes | Yes |
wan-3.0 | Yes | Yes |
minimax-h3 and minimax-h3-max | Yes | Yes |
gemini-omni-flash-1.1 | Yes | No |
higgsfield-genjutsu | Yes | No |
h3-max-recast | Yes | No |
What rules apply when I send audio?
Send public HTTPS URLs only. On the Video 1.0 and Auto shape, reference_audio_urls takes 1 to 3 URLs and requires at least one reference image or video, so audio alone is not a valid request there.
On /v1/videos, put audio in input_references with type set to the audio type the catalog lists, and read supported_input_references first. Do not assume that a reference means the clip's soundtrack will be that audio; test one short clip and listen before you batch.
How do I check a model before submitting?
Fetch the catalog entry and test for audio_url in the list. The check is two lines, and it saves a rejected request and a debugging round trip.
import os
import requests
r = requests.get(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
if "audio_url" in (m.get("supported_input_references") or []):
print(m["id"], "accepts audio references")What if I need a specific voice track on the clip?
If the goal is a particular narration or music bed, a reference is the wrong tool. Generate the clip, then place your audio on it with the timeline tools, which keep the video and swap the sound under your control. Use reference audio when you want the model to take the sound into account while it generates.
Sources
Related posts
More in Models
- Same voice across AI video clips: Kling Omni voice vs Seedance audio
Kling 3.0 Omni binds a voice from a 5-30 second sample; Seedance takes up to 3 audio references. What each page says and what Sume lets you send.
- Seed Audio 1.0 on OpenRouter: text to speech, and what Sume offers
ByteDance's Seed Audio 1.0 is listed on OpenRouter as text-to-speech with raw MP3 or PCM output. Sume's TTS Router is Cartesia Sonic only today.
- Seedance 2.5 concert prompt uses 18 images: fitting Sume's 9
Seed's Seedance 2.5 concert example uses 18 reference images. Sume's Video Router takes at most 9 for seedance-2.5, so merge groups into sheets.
- Seedance 2.5 aspect ratios on Sume: eight values incl. adaptive
Sume lists eight aspect ratio values for seedance-2.5: auto, adaptive, 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16. Uses in ad placements and what is missing.
Written by Sume