AI Video With Built-In Audio: Which Models Always Render Sound
Which video models always render audio, which have a switch, and which have none: Veo 3.1, Gemini Omni, Kling 3, H3, Grok and Seedance on Sume.

Native audio changes cost, editing and licensing, so whether a model lets you turn it off is a real selection criterion. As of 2026-10-03 the catalog splits three ways: always on, optional, and not offered.
Always on
Google's Veo page says native audio is always on for Veo 3.1. On Sume, gemini-omni-flash-1.1 rejects generate_audio: false, and minimax-h3 renders native stereo audio on every clip. minimax-h3-max also lists native audio.
If your edit will replace the soundtrack, an always-on model still produces audio you pay for and then discard. That is usually small, but it is not free.
Optional or absent
Kling is the one with a price switch: Sume lists kling-3 at $0.14 per second with audio off and $0.21 with audio on, and Kling's own page describes multi-shot generation with native audio and lip sync. Seedance generates audio and video jointly per ByteDance, and Sume's generate_audio field defaults to the model's capability. Grok Imagine Video 1.5 on Sume has no audio toggle at all.
| Model | Audio | Control |
|---|---|---|
| gemini-omni-flash-1.1 | Native, synced | Always on; false is rejected |
| minimax-h3 | Native stereo | Always on |
| minimax-h3-max | Native audio | Listed as native |
| kling-3 | Optional | $0.14 off, $0.21 on per second |
| seedance-2.5 and 2.0 family | Joint audio-video | generate_audio defaults to model capability |
| grok-imagine-video-1.5 | None toggled | No audio control |
What audio does not do
Sume's docs are explicit that video models do not lip-sync to generated TTS or to a later voice-over. Native audio is generated with the picture, so it can include ambience and sound effects, but a specific script spoken by a specific face belongs on a lip-sync route. See the lip-sync decision.
Practical rule: for wordless B-roll, use native audio and keep it as a scratch track; for dialogue, lock the voice first and sync the face to it.
Sources
Related posts
More in Models
- AI voice agent latency budget: who owns which 100 milliseconds
Vendor numbers for speech-to-text, the language model and text-to-speech side by side, with a note on where an async file API like Sume belongs and where not.
- Arabic text to speech API: send language ar to Sume TTS
Arabic is on Cartesia's Sonic 3.6, 3.5 and 3 but never on Sonic 2. How to send ar through Sume TTS, which voice id to use, and what an Arabic script costs.
- AudioCraft weights are CC-BY-NC: which models that covers
AudioCraft's code is MIT but its README puts the model weights under CC-BY-NC 4.0, including MusicGen and AudioGen. What that means for ad and video audio.
- AuK: MIT speech model that edits audio, and what Sume covers
AuK is Tencent's 1.5B MIT-licensed speech model for generation and editing. What its card lists, what it omits, and which parts Sume's audio tools cover.
Written by Sume