generate_audio false: which Sume video models let you turn sound off
Omni and MiniMax H3 and H3 Max always make audio and reject generate_audio false. Wan 3.0, Seedance 2.5 and Kling 3 take the toggle; Kling prices sound apart.

Only some Sume video models let you turn sound off. Gemini Omni Flash 1.1, MiniMax H3 and MiniMax H3 Max always produce audio, and generate_audio: false returns 400. Wan 3.0 and Seedance 2.5 take an optional generate_audio. Kling 3 takes the toggle too, and its list price differs with sound on or off.
What the model cards say
The Video Router model cards on main state it per model. Omni: always-on native synced audio, so omit the field. MiniMax H3 and H3 Max: native stereo audio, always on. Wan 3.0 and Seedance 2.5: audio is optional through generate_audio. sume/auto follows Omni: sending generate_audio: false fails with a 400 instead of silently routing elsewhere.
| Model id | Sound | generate_audio false |
|---|---|---|
| gemini-omni-flash-1.1 | Always on, synced | 400 |
| sume/auto | Always on | 400, never rerouted to Seedance |
| minimax-h3 | Always on, stereo | Rejected |
| minimax-h3-max | Always on, stereo | Rejected |
| wan-3.0 | Optional | Accepted |
| seedance-2.5 | Optional | Accepted |
| kling-3 | Optional | Accepted |
| grok-imagine-video-1.5 | No toggle in the schema | n/a |
Price of sound
For Wan 3.0 and Omni the repo has one per-second rate, so sound does not change the price. Seedance bills on video tokens only. Kling Video v3 Pro is the exception: the pricing comment in pricing-tables.ts records a list of $0.112 per second with audio off and $0.168 with audio on. At the house x 1.25 that is $0.14 and $0.21 per second, so a 10 s Kling clip is $1.40 without sound and $2.10 with it. Check the current catalog before you rely on those numbers.
If you need silence
The generation call has no mute flag for the always-on models. The lever you have is the prompt: ask for no dialogue and quiet room tone only. If a hard-silent file matters, choose a model from the toggle rows (Wan 3.0, Seedance 2.5 or Kling 3) and send generate_audio: false.
When you assemble clips in Timeline 1.0, the soundtrack comes from the audio spine you declare. Check the output of a test render with a sound meter before you publish.
Sending the right body
Omit the field for always-on models. Send it only where the toggle exists.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"wan-3.0","prompt":"A product turntable, silent studio","duration":5,"resolution":"720p","generate_audio":false}'Checking from the catalog
The generate_audio field on each model row in GET /v1/videos/models tells you whether the model can make audio at all. Whether the toggle exists is described in the model's notes, and the API enforces it with a 400. Do a one-line test with the shortest duration if a doc and your result disagree.
Remember that the default for the models with a toggle is on, per the user guide: the default is the audio capability of the model. If you want silence, send false explicitly.
Related posts
More in Models
- GPT Image 2.5 above 2560x1440 is experimental: a safe size ladder
OpenAI marks sizes over 2560x1440 experimental on GPT Image models. A ladder of valid sizes up to 3840x2160 and a Python check for Sume's image_size.
- GPT Image 2.5 quality: OpenAI defaults to auto, Sume to high
OpenAI's default quality for GPT Image 2.5 is auto; Sume's is high when omitted. Why that differs, what auto reserves on Sume, and how to pin quality.
- gpt-live-transcribe: realtime STT at $0.017 a minute
gpt-live-transcribe costs $0.017 a minute ($1.02 an hour) and runs only on the Realtime transcription sessions endpoint. Use gpt-transcribe for files.
- xAI recommends Grok Imagine Video 1.5 ($0.08/s) over the $0.05 model
xAI lists grok-imagine-video-1.5 at $0.080 per second and grok-imagine-video at $0.050, and recommends 1.5. Sume carries only 1.5, image-to-video.
Written by Sume