AI video audio by model: toggle, always on, or none on Sume
Seedance, Wan 3.0 and Kling 3 let you switch audio; MiniMax H3 and Gemini Omni Flash always make it; Grok Imagine has none. What each row does.

On Sume, audio depends on the model row. Seedance 2.5, the Seedance 2.0 rows, Wan 3.0 and Kling 3 generate audio when generate_audio is on; MiniMax H3, MiniMax H3 Max and Gemini Omni Flash 1.1 always produce audio and reject or ignore the switch; Grok Imagine Video 1.5 and Genjutsu produce none. Decide this before you choose a model, because it changes both the price and your edit.
Audio behaviour per row
The catalog constraints say it directly. For the always-on rows, omit the field; Gemini Omni Flash rejects generate_audio: false outright.
| Model id | Audio |
|---|---|
| seedance-2.5 and the Seedance 2.0 rows | optional generate_audio |
| wan-3.0 | optional generate_audio |
| kling-3 | optional; separate audio-on and audio-off per-second prices |
| minimax-h3, minimax-h3-max | native stereo, always on, no toggle |
| gemini-omni-flash-1.1 | native synced audio, always on; false is rejected |
| grok-imagine-video-1.5 | none |
| h3-max-recast | keeps the source soundtrack |
Why it matters for editing
If you plan to lay a music bed or a voice-over, an always-on row hands you a clip with its own sound that you then have to detach or duck. A toggle row lets you generate silent and mix cleanly. If you want synced ambient sound or speech from the model, choose a row that makes it by default and check the transcript with a video inspect step before you ship.
Cost
Kling 3 is the one row with two published per-second prices, one with audio and one without. For the others, audio is part of the model's price. Read the pricing fields in GET /v1/video-router/models for the current figures.
Sources
Related posts
More in Models
- How to choose an AI video model: five questions before you pin one
Length, resolution, audio, start image and references decide the model, not the leaderboard. Five questions mapped to Sume's video catalog.
- Which AI video models can't do text-to-video? Rows that need a source
Grok Imagine Video 1.5, Genjutsu Motion Transfer and H3 Max Recast refuse a prompt-only request. What each needs, and which rows accept text alone.
- AI video reference limits: how many images and clips per model
Reference image and reference video caps for Wan 3.0, MiniMax H3, Gemini Omni Flash, Genjutsu and H3 Max Recast on Sume, in one table with the odd limits.
- AI voice agent latency budget: who owns which 100 milliseconds
Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.
Written by Sume