AI video with audio: which models always add sound, which are optional

On Sume, minimax-h3, minimax-h3-max and gemini-omni-flash-1.1 always add native audio; kling-3, Seedance and Wan make it optional; grok-imagine-video-1.5 none.

4 min readSume
All posts

Sume's video models split three ways on sound: always on (minimax-h3, minimax-h3-max, gemini-omni-flash-1.1), optional through generate_audio (kling-3, seedance-2.5, the Seedance 2 family, wan-3.0) and none (grok-imagine-video-1.5). When you leave generate_audio out, it follows the model's own capability, so an optional model has sound on by default.

From Video generation and catalog code, 2026-09-29.

Which model does what?

Audio behavior per Sume video model from the docs and catalog code, read 2026-09-29.
ModelAudio
minimax-h3, minimax-h3-maxAlways on, native stereo
gemini-omni-flash-1.1Always on, native synced
kling-3Optional, default on; $0.14 vs $0.21 per second
Seedance 2.5 and 2.xOptional
wan-3.0Optional
grok-imagine-video-1.5None

How do I turn sound off?

Send generate_audio: false on models that allow it. On always-on models the request is not the lever; pick another model.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: silent-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "A paper boat drifting down a rain gutter.",
    "duration": 5,
    "generate_audio": false
  }'

Can I supply my own audio?

Only where supported_input_references lists audio_url: the Seedance 2.x models, wan-3.0, minimax-h3 and minimax-h3-max. gemini-omni-flash-1.1 and kling-3 list no audio input.

Sources

Related posts

More in Models

All Models posts

Written by Sume