generate_audio false on Gemini Omni: accepted, clip keeps its audio
Video Router accepts generate_audio false on gemini-omni-flash-1.1 and ignores it; MiniMax H3 answers 400. Check the model before you rely on a mute flag.

If you send generate_audio: false to gemini-omni-flash-1.1 on Video Router, the request is accepted and the clip still carries native synced audio. For minimax-h3 and minimax-h3-max the same flag is a 400: "model minimax-h3 always produces native stereo audio; omit generate_audio." For grok-imagine-video-1.5 it is refused because that model takes no such field.
Docs and code differ on Gemini Omni
The Video Router docs page says native audio is always on and that the API rejects generate_audio: false. The validation code on main does something different: it accepts the flag as a no-op. A code comment explains why. A caller that pinned both the model and generate_audio: false for b-roll could not satisfy a rejection, and it caused many wasted create calls, so the flag is accepted and has no effect. The schema tests assert both true and false parse for this model. Trust behaviour you have observed over a sentence on a docs page, and read the code or test when it matters.
| Model | generate_audio false | What you get |
|---|---|---|
| gemini-omni-flash-1.1 | Accepted, no effect | Native synced audio, always |
| minimax-h3, minimax-h3-max | 400: always produces native stereo audio; omit generate_audio | Native stereo audio, always |
| grok-imagine-video-1.5 | 400: generate_audio is not supported by model grok-imagine-video-1.5. | No audio field exists |
| Seedance, Wan, kling-3 | No refusal in the schema | Read the model entry for the audio default |
How to get a silent result
If you need picture without sound for an edit, do not rely on the flag for the always-on models. Pick a model whose catalog entry offers an audio toggle, and verify the result: video_inspect on a stored clip returns probe.has_audio, and a frames: false inspect is enough to read it.
Guard in the client
The snippet encodes the three behaviours so a pipeline stops asking for something the model will ignore or refuse. It prints a decision and sends nothing.
ALWAYS_ON_IGNORED = {"gemini-omni-flash-1.1"}
ALWAYS_ON_REFUSED = {"minimax-h3", "minimax-h3-max"}
NO_FIELD = {"grok-imagine-video-1.5"}
def audio_plan(model: str, want_audio: bool) -> str:
if want_audio:
return "omit the flag; sound is on" if model in ALWAYS_ON_REFUSED | ALWAYS_ON_IGNORED else "send generate_audio true"
if model in ALWAYS_ON_REFUSED | NO_FIELD:
return "do not send the flag; choose another model"
if model in ALWAYS_ON_IGNORED:
return "flag is ignored; the clip keeps its audio"
return "send generate_audio false"
for m in ("gemini-omni-flash-1.1", "minimax-h3", "kling-3"):
print(m, "->", audio_plan(m, False))
When not to trust a silent flag
Always listen to the output, or probe its streams, before you build a timeline on the assumption that a clip is mute. A no-op flag is worse than a loud error because nothing tells you it did nothing. The same goes for the Gemini Omni reference rules: it takes images and short videos but no audio references, and aspect ratio is limited to 16:9 or 9:16.
If your pipeline reads a model's capabilities at start-up, treat "audio: true with no toggle" as a distinct state from "audio: true with a toggle". The first means the clip will have sound whatever you send, which affects mixing, loudness and licensing checks downstream.
Sources
Related posts
More in Models
- GPT Image 2.5 above 2560x1440 is experimental: a safe size ladder
OpenAI marks sizes over 2560x1440 experimental on GPT Image models. A ladder of valid sizes up to 3840x2160 and a Python check for Sume's image_size.
- GPT Image 2.5 quality: OpenAI defaults to auto, Sume to high
OpenAI's default quality for GPT Image 2.5 is auto; Sume's is high when omitted. Why that differs, what auto reserves on Sume, and how to pin quality.
- gpt-live-transcribe: realtime STT at $0.017 a minute
gpt-live-transcribe costs $0.017 a minute ($1.02 an hour) and runs only on the Realtime transcription sessions endpoint. Use gpt-transcribe for files.
- xAI recommends Grok Imagine Video 1.5 ($0.08/s) over the $0.05 model
xAI lists grok-imagine-video-1.5 at $0.080 per second and grok-imagine-video at $0.050, and recommends 1.5. Sume carries only 1.5, image-to-video.
Written by Sume