Silent AI video: generate_audio false or drop the audio after
Seedance, Kling and Wan take generate_audio false; Omni and MiniMax H3 always make sound. How to get a silent clip on Sume, and what it costs.

To get a silent AI video from Sume, set generate_audio: false on a model that lets you, and strip the audio afterwards on the models that do not. Seedance 2.x and 2.5, Kling 3 and Wan 3.0 take generate_audio. Gemini Omni Flash 1.1, MiniMax H3 and MiniMax H3 Max always produce native audio and you omit the flag, so for those you run the finished clip through Sume's video-trim with audio: "drop". Grok Imagine Video 1.5 has no audio at all.
The audio behavior comes from Sume's Video Router docs, Video generation docs and video trim docs, read on 2026-10-03.
Which models can generate without audio?
When you leave generate_audio out, it defaults to the model's own audio capability, so a clip may come back with sound you did not ask for. If you want silence, say so in the request on the models that allow it.
Sending the flag on the wrong model is an error rather than a no-op. Omni rejects generate_audio: false, and Grok Imagine Video 1.5 rejects generate_audio entirely, so a template that always sends the flag will fail on those ids.
| Model id | Audio on output | What to send |
|---|---|---|
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini | Optional | generate_audio: false for silence |
kling-3 | Optional, with separate audio-on and audio-off price rows | generate_audio: false for silence |
wan-3.0 | Optional | generate_audio: false for silence |
gemini-omni-flash-1.1 | Always on | Omit the flag; false is rejected |
minimax-h3, minimax-h3-max | Always on | Omit the flag |
grok-imagine-video-1.5 | None | Omit the flag; it is rejected |
How do I mute a clip after it is generated?
Use video trim. It takes one clip already hosted on media.sume.com, which is what a finished job's artifact URL is, and cuts a [start, end) range into a new MP4. Its audio option is keep (the default) or drop. To mute without shortening, set start to 0 and duration to the clip's length.
The docs price a trim at $0.02 per job, with no provider inference, and note that precision: "exact" re-encodes the video while keyframe copies the stream. Muting with exact is the safe choice because the cut lands where you asked.
curl -X POST https://api.sume.com/v1/video-trim \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: mute-clip-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4",
"start": 0,
"duration": 8,
"audio": "drop"
}'Does a silent clip cause problems later?
Sometimes. Sume's timeline compose docs treat a mute video as a warning, compose_video_has_no_audio, not a failure: the clip still renders, and the Timeline 1.0 spine supplies the audio at assemble time. That is the design when you plan a voiceover or music bed over generated footage.
Captions are the other place silence matters. A caption job needs speech to align, so a silent clip has nothing to transcribe; the stored post MiniMax H3 clip, add captions: a silent clip fails covers the fix, which is to provide cue text yourself.
- Generate silent on Seedance, Kling or Wan when the soundtrack is decided elsewhere.
- Drop audio with video trim on Omni and MiniMax outputs.
- Keep the native audio when the model's dialogue or ambience is the point, and extract it with audio detach if you want it as a separate file.
- Never send
generate_audiofrom a shared template without checking the model id.
Which route is cheaper?
Asking for no audio on a model with optional audio avoids a step and, on Kling 3, uses the audio-off price row. Muting afterwards adds a $0.02 trim job to a clip you already paid to generate with audio. Read the live prices from GET /v1/videos/models or GET /v1/catalog, because the audio rows differ by model and the catalog is the source of truth. The related post AI video with sound: generate_audio covers the audio-on side.
Sources
Related posts
More in Models
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
- Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume