AI video models with sound: which toggle it, which are always on

Seedance 2.5, Wan 3.0, Kling 3 and MiniMax H3 all return audio on Sume. Kling prices audio separately; H3 and Omni cannot turn it off. What each does.

5 min readSume
All posts

Every video model on Sume's trend list makes sound, but they differ in whether you can turn it off. kling-3 has an audio-on rate and an audio-off rate. minimax-h3 and minimax-h3-max always produce audio and reject generate_audio: false with a 400, and gemini-omni-flash-1.1 always produces synced audio and ignores the flag. seedance-2.5 and wan-3.0 report audio support in the catalog.

Magic Hour's tracker (read 2026-10-06) agrees on the direction: Seedance 2.5 is listed as audio-video, MiniMax H3 as native stereo, Kling VIDEO 3.0 as native audio. The Sume side is in the video generation docs and the Video Router docs. See the tracker for the vendor wording.

Audio behavior by id

The generate_audio flag in GET /v1/videos/models tells you whether a model can make audio. The constraints text in the router catalog tells you whether it is a toggle.

Sume catalog audio behavior; Kling price is the billed rate (list x 1.25), read 2026-10-06.
Sume idAudioCan you turn it off?Audio-related price
kling-3yesyes, priced per second$0.14 off, $0.21 on, per second
minimax-h3native stereono toggleincluded in the per-second rate
minimax-h3-maxnative stereono toggleincluded in the per-second rate
gemini-omni-flash-1.1native syncedno; false has no effectincluded
seedance-2.5yessee generate_audio in the catalogper video token, no audio line
wan-3.0yessee generate_audio in the catalogper second by resolution

What a silent request does on an always-on model

If you send generate_audio: false to minimax-h3 or minimax-h3-max, the request fails with a 400 that tells you to omit the field. On gemini-omni-flash-1.1 the flag is accepted but changes nothing: the clip still carries audio, so do not read the absence of an error as silence. Either way, there is no price saving on the always-on models, because audio is included in the per-second rate.

On Kling the choice is real money. A 10-second clip is $1.40 silent and $2.10 with audio, so a batch of 100 silent B-roll clips saves $70.00 against leaving sound on.

Read the flag instead of assuming

Audio support is a catalog field, so a script can decide before it submits:

import os, requests

r = requests.get(
    "https://api.sume.com/v1/videos/models",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
    if m["id"] in ("kling-3", "minimax-h3", "seedance-2.5", "wan-3.0"):
        print(m["id"], "generate_audio:", m["generate_audio"],
              "refs:", m["supported_input_references"])

Where to read more

The H3 detail is in MiniMax H3 native stereo audio, the Kling price split in Kling 3.0 audio on or off, and the general limits in the Kling 3.0 API post. Before a batch, validate each request against the catalog.

Sound is not the same as speech

Native audio in these models means a generated soundtrack: ambience, effects and often speech that fits the picture. It does not mean a script you control. If a line of dialogue must be word-perfect, generate the voice separately with text to speech, and use a model that takes an audio reference or a lip-sync step. seedance-2.5, wan-3.0, minimax-h3 and minimax-h3-max all accept an audio_url reference in input_references, and kling-3 does not take any reference inputs.

One more check is worth doing before a paid batch: listen to a short sample. Always-on audio on H3 is stereo, which matters if the clip goes to a platform that downmixes. A silent Kling clip is the cheapest way to get B-roll that you will score with music afterwards, and the difference over a batch is the $70.00 computed above.

What to do with a clip whose sound you do not want

On the always-on models you cannot ask for silence, but you can remove it afterwards. Take the picture and put your own music or voice under it with a Timeline 1.0 render, whose audio spine is a file you choose. That costs a render at $0.10 per output minute instead of a new generation, and it keeps the picture you paid for. Check the first render to confirm the original sound is not carried through. If the sound itself is the product, a short test at 480p shows you whether the model's ambience and effects suit the scene before you commit to 1080p.

Sources

Related posts

More in Models

All Models posts

Written by Sume