AI video models with sound: which toggle it, which are always on
Seedance 2.5, Wan 3.0, Kling 3 and MiniMax H3 all return audio on Sume. Kling prices audio separately; H3 and Omni cannot turn it off. What each does.

Every video model on Sume's trend list makes sound, but they differ in whether you can turn it off. kling-3 has an audio-on rate and an audio-off rate. minimax-h3 and minimax-h3-max always produce audio and reject generate_audio: false with a 400, and gemini-omni-flash-1.1 always produces synced audio and ignores the flag. seedance-2.5 and wan-3.0 report audio support in the catalog.
Magic Hour's tracker (read 2026-10-06) agrees on the direction: Seedance 2.5 is listed as audio-video, MiniMax H3 as native stereo, Kling VIDEO 3.0 as native audio. The Sume side is in the video generation docs and the Video Router docs. See the tracker for the vendor wording.
Audio behavior by id
The generate_audio flag in GET /v1/videos/models tells you whether a model can make audio. The constraints text in the router catalog tells you whether it is a toggle.
| Sume id | Audio | Can you turn it off? | Audio-related price |
|---|---|---|---|
kling-3 | yes | yes, priced per second | $0.14 off, $0.21 on, per second |
minimax-h3 | native stereo | no toggle | included in the per-second rate |
minimax-h3-max | native stereo | no toggle | included in the per-second rate |
gemini-omni-flash-1.1 | native synced | no; false has no effect | included |
seedance-2.5 | yes | see generate_audio in the catalog | per video token, no audio line |
wan-3.0 | yes | see generate_audio in the catalog | per second by resolution |
What a silent request does on an always-on model
If you send generate_audio: false to minimax-h3 or minimax-h3-max, the request fails with a 400 that tells you to omit the field. On gemini-omni-flash-1.1 the flag is accepted but changes nothing: the clip still carries audio, so do not read the absence of an error as silence. Either way, there is no price saving on the always-on models, because audio is included in the per-second rate.
On Kling the choice is real money. A 10-second clip is $1.40 silent and $2.10 with audio, so a batch of 100 silent B-roll clips saves $70.00 against leaving sound on.
Read the flag instead of assuming
Audio support is a catalog field, so a script can decide before it submits:
import os, requests
r = requests.get(
"https://api.sume.com/v1/videos/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
for m in r.json()["data"]:
if m["id"] in ("kling-3", "minimax-h3", "seedance-2.5", "wan-3.0"):
print(m["id"], "generate_audio:", m["generate_audio"],
"refs:", m["supported_input_references"])Where to read more
The H3 detail is in MiniMax H3 native stereo audio, the Kling price split in Kling 3.0 audio on or off, and the general limits in the Kling 3.0 API post. Before a batch, validate each request against the catalog.
Sound is not the same as speech
Native audio in these models means a generated soundtrack: ambience, effects and often speech that fits the picture. It does not mean a script you control. If a line of dialogue must be word-perfect, generate the voice separately with text to speech, and use a model that takes an audio reference or a lip-sync step. seedance-2.5, wan-3.0, minimax-h3 and minimax-h3-max all accept an audio_url reference in input_references, and kling-3 does not take any reference inputs.
One more check is worth doing before a paid batch: listen to a short sample. Always-on audio on H3 is stereo, which matters if the clip goes to a platform that downmixes. A silent Kling clip is the cheapest way to get B-roll that you will score with music afterwards, and the difference over a batch is the $70.00 computed above.
What to do with a clip whose sound you do not want
On the always-on models you cannot ask for silence, but you can remove it afterwards. Take the picture and put your own music or voice under it with a Timeline 1.0 render, whose audio spine is a file you choose. That costs a render at $0.10 per output minute instead of a new generation, and it keeps the picture you paid for. Check the first render to confirm the original sound is not carried through. If the sound itself is the product, a short test at 480p shows you whether the model's ambience and effects suit the scene before you commit to 1080p.
Sources
Related posts
More in Models
- Which new video model for which job: Seedance, Wan, Kling, H3, Omni
A decision table for the five new video models on Sume, by clip length, resolution, sound and price per second. Veo 3.1 and LTX-2.5 are not on Sume.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
Written by Sume