AI Video With Built-In Audio: Which Models Always Render Sound

Which video models always render audio, which have a switch, and which have none: Veo 3.1, Gemini Omni, Kling 3, H3, Grok and Seedance on Sume.

5 min readSume
All posts

Native audio changes cost, editing and licensing, so whether a model lets you turn it off is a real selection criterion. As of 2026-10-03 the catalog splits three ways: always on, optional, and not offered.

Always on

Google's Veo page says native audio is always on for Veo 3.1. On Sume, gemini-omni-flash-1.1 rejects generate_audio: false, and minimax-h3 renders native stereo audio on every clip. minimax-h3-max also lists native audio.

If your edit will replace the soundtrack, an always-on model still produces audio you pay for and then discard. That is usually small, but it is not free.

Optional or absent

Kling is the one with a price switch: Sume lists kling-3 at $0.14 per second with audio off and $0.21 with audio on, and Kling's own page describes multi-shot generation with native audio and lip sync. Seedance generates audio and video jointly per ByteDance, and Sume's generate_audio field defaults to the model's capability. Grok Imagine Video 1.5 on Sume has no audio toggle at all.

Audio behavior on Sume, read 2026-10-03
ModelAudioControl
gemini-omni-flash-1.1Native, syncedAlways on; false is rejected
minimax-h3Native stereoAlways on
minimax-h3-maxNative audioListed as native
kling-3Optional$0.14 off, $0.21 on per second
seedance-2.5 and 2.0 familyJoint audio-videogenerate_audio defaults to model capability
grok-imagine-video-1.5None toggledNo audio control

What audio does not do

Sume's docs are explicit that video models do not lip-sync to generated TTS or to a later voice-over. Native audio is generated with the picture, so it can include ambience and sound effects, but a specific script spoken by a specific face belongs on a lip-sync route. See the lip-sync decision.

Practical rule: for wordless B-roll, use native audio and keep it as a scratch track; for dialogue, lock the voice first and sync the face to it.

Sources

Related posts

More in Models

All Models posts

Written by Sume