Wan 3.0 sound toggle vs Sume generate_audio for Wan

Wan 3.0 lists a sound toggle at QwenCloud. On Sume, generate_audio defaults to the model's audio capability and some models reject false.

4 min readSume
All posts

QwenCloud lists a sound toggle among Wan 3.0 features. On Sume, generate_audio is the closest field: it defaults to the model's audio capability, and whether false is allowed depends on the model. For gemini-omni-flash-1.1, generate_audio: false is rejected, and the minimax-h3 rows list native stereo audio with no toggle.

The Kling counterpart is in Kling 4.0 stereo audio vs the generate_audio flag.

What did QwenCloud say about sound?

The changelog lists audio generation and a sound toggle as Wan 3.0 features. It does not say what the toggle sends on the wire.

What does generate_audio do on Sume?

The Video API docs describe generate_audio as whether to generate audio alongside the video, defaulting to the model's audio capability. The model listing's generate_audio field reports whether a model can generate an audio track. The Wan row in the Video Router docs lists audio.

Audio behavior by row, read 2026-10-01.
Model rowWhat the docs say
wan-3.0Lists audio
minimax-h3Native stereo audio (no toggle)
minimax-h3-maxNative stereo audio (no toggle)
gemini-omni-flash-1.1Native audio always on; generate_audio: false is rejected

Can I turn audio off for Wan?

The docs reviewed here say the Wan row supports audio, and that the field defaults to the model's capability. They do not state a rule for sending generate_audio: false to wan-3.0, so test it on a short clip and read the response before relying on it.

How do I check a model's audio support?

Call GET /v1/video-router/models and read generate_audio for the model. If a model has no toggle, mute the audio in your editor or join step instead.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume