Reference audio for AI video: which Sume models take audio_url

Seedance 2.x, Wan 3.0 and MiniMax H3 honor audio references on Sume; Gemini Omni, Genjutsu and Recast take video but not audio. Rules and a check.

5 min readSume
All posts

On Sume, audio references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. Check supported_input_references for audio_url before you send one, because a model that does not list it rejects the field.

The question comes up with every launch that mentions audio. Kling's announcement for Kling 4.0 (read 2026-10-02) describes stereo audio with improved lip-sync, but that is audio the model produces, not audio you supply. Sume's docs treat supplied audio as a reference input, and that support is per model.

Which models honor an audio reference?

The table lists what Sume's video generation and Video Router docs state. Your live catalog is the final word.

Reference support by model as stated in Sume docs, read 2026-10-02
ModelVideo referenceAudio reference
Seedance 2.xYesYes
wan-3.0YesYes
minimax-h3 and minimax-h3-maxYesYes
gemini-omni-flash-1.1YesNo
higgsfield-genjutsuYesNo
h3-max-recastYesNo

What rules apply when I send audio?

Send public HTTPS URLs only. On the Video 1.0 and Auto shape, reference_audio_urls takes 1 to 3 URLs and requires at least one reference image or video, so audio alone is not a valid request there.

On /v1/videos, put audio in input_references with type set to the audio type the catalog lists, and read supported_input_references first. Do not assume that a reference means the clip's soundtrack will be that audio; test one short clip and listen before you batch.

How do I check a model before submitting?

Fetch the catalog entry and test for audio_url in the list. The check is two lines, and it saves a rejected request and a debugging round trip.

import os
import requests

r = requests.get(
    "https://api.sume.com/v1/videos/models",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
    if "audio_url" in (m.get("supported_input_references") or []):
        print(m["id"], "accepts audio references")

What if I need a specific voice track on the clip?

If the goal is a particular narration or music bed, a reference is the wrong tool. Generate the clip, then place your audio on it with the timeline tools, which keep the video and swap the sound under your control. Use reference audio when you want the model to take the sound into account while it generates.

Sources

Related posts

More in Models

All Models posts

Written by Sume