kling-3 on Sume: generate_audio true or false, and what changes
kling-3 has an audio toggle; MiniMax H3 and Omni audio is always on. The catalog lists separate audio-on and audio-off list rates.

On Sume generate_audio is optional and defaults to the model's audio capability, so a plain kling-3 request produces audio. Send generate_audio: false for a silent clip. The catalog lists two list rates for kling-3, one with audio on and one with audio off, so the flag is also a cost choice. For MiniMax H3 and H3 Max and for Gemini Omni Flash 1.1 the audio is always on and generate_audio: false is rejected with a 400.
Kling's page describes Kling 3.0 audio as single-channel and Kling 4.0 as two-channel stereo (kling.ai, read 2026-10-03); Kling 4.0 is not callable on Sume. The Sume behaviour is in the Video generation docs and the Video Router docs.
Which ids let you turn audio off
Sume does not print a single audio table in the docs, so the safest read is the live model list: each entry has generate_audio: true when the model can produce audio. Beyond that flag, three ids reject generate_audio: false outright because their audio is always produced.
If you request audio from a model that cannot make it, the answer is also a 400 ('does not support generate_audio'), not a silent clip.
| Sume id | Audio | generate_audio false |
|---|---|---|
| kling-3 | optional | accepted, silent clip |
| minimax-h3 / minimax-h3-max | native stereo, always on | 400, omit the field |
| gemini-omni-flash-1.1 | native synced, always on | 400, omit the field |
| seedance-2.5 and other audio-capable ids | model capability | read generate_audio in the model list |
When to turn it off
Turn audio off when you will lay your own voice or music over the clip: an avatar voice-over, a licensed track, or a captioned social cut. Generated sound would only be muted later, and with kling-3 you avoid paying for it.
Leave it on for dialogue or sound-effect clips, and check the result: the video tools can inspect the file, and the audio track tells you whether the model produced speech or ambience.
Do not assume the price
Sume bills list times 1.25 and the list price is shown per model in the models endpoint. This post does not quote a figure because list prices change; read pricing for kling-3 at request time and compute your own total for the number of seconds you will generate.
A worked habit: store the list rate with each job, so later you can explain a bill without guessing which flag was set.
Sources
Related posts
More in Developers
- Kling 3 input_references 400 unsupported_capability: fix
Sending input_references to kling-3 on Sume returns 400 unsupported_capability. Move the image to frame_images or pick a model that takes references.
- Kling 3.0 Motion Control on Sume: four routes, one stored model id
POST /v1/kling/3.0/motion-control and three aliases all store kling/3.0/motion-control. How to pick a route, what the body needs, and what the job reserves.
- tts_language_script_mismatch: Korean TTS needs a Hangul syllable
Sume TTS returns 400 tts_language_script_mismatch when language is ko but the text has no Hangul syllable. Usual cause: UTF-8 decoded as Latin-1. Fix it.
- Kotlin 2.4.20 allDistinctBy: unique Idempotency-Keys per batch
Kotlin 2.4.20 adds allDistinctBy. Use it to assert one Idempotency-Key per intent before a batch of Sume image submits, then post with java.net.http.
Written by Sume