'reference_audio_urls requires a reference image or video'

Audio cannot be the only reference on Sume video models. Add one image or video reference. Which models take audio references, and their caps.

4 min readSume
All posts

Video Router refuses a request that has reference_audio_urls but no reference_image_urls or reference_video_urls, with the message "reference_audio_urls requires at least one reference image or video." An audio file can steer a clip, but it cannot be the only reference. Add at least one image or video.

Which models take audio references at all

The check is general, but only some models accept audio references in the first place. The catalog marks audio references as true for the Seedance 2.x family, wan-3.0, minimax-h3 and minimax-h3-max. It marks them false for gemini-omni-flash-1.1, grok-imagine-video-1.5, kling-3, higgsfield-genjutsu and h3-max-recast. Gemini Omni gives a model-specific refusal: "reference_audio_urls is not supported by model gemini-omni-flash-1.1."

Audio reference support and caps in the Video Router catalog and schema on main (read 2026-10-05)
ModelAudio referencesCap in schema or catalog
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-miniYesAt most 3 files
wan-3.0YesAt most 5 files, 15 s total (catalog note)
minimax-h3, minimax-h3-maxYesAt most 3 files, 2 to 15 s each, 15 s combined (catalog note)
gemini-omni-flash-1.1NoRefused by name
kling-3, grok-imagine-video-1.5NoReference lists refused entirely

What counts toward the limits

The MiniMax H3 notes add that images, videos and audios together may not exceed 12 files, and the general schema caps the total at 12 as well for models outside Wan and Gemini. So a request with 9 images, 3 videos and 3 audios is over the line before any single list is. The neighbouring post on the 12-file total lists every cap.

A safe request

The body below sends one reference image and one audio file to Wan 3.0. It builds the dictionary and checks the rule locally, so it prints nothing secret and creates no job. Replace the URLs with public HTTPS files you own.

def check_audio_refs(body: dict) -> None:
    audio = body.get("reference_audio_urls")
    if audio and not (body.get("reference_image_urls") or body.get("reference_video_urls")):
        raise ValueError("audio needs at least one image or video reference")

body = {
    "model": "wan-3.0",
    "prompt": "The singer turns to camera on the last chorus",
    "resolution": "720p",
    "duration": 10,
    "reference_image_urls": ["https://example.com/singer.png"],
    "reference_audio_urls": ["https://example.com/chorus.mp3"],
}
check_audio_refs(body)
print(sorted(body))

When audio is the wrong lever

If all you have is a voice track and no picture, use a model that makes the picture from text and describe the sound in the prompt, or generate a still first. Audio references are guidance, not a way to turn a recording into a talking clip by themselves.

Also remember that the rule is about presence, not usefulness. A throwaway image added only to satisfy the check will still influence the clip. Pick an image that really belongs to the scene, such as the subject or location the audio is about, so the extra reference helps instead of distorting the result. Keep the audio short and clean, inside the catalog limits listed above, and test one clip before you queue a batch of variations.

Sources

Related posts

More in Models

All Models posts

Written by Sume