Silent clips after Sora: which Sume models take generate_audio false

Gemini Omni Flash 1.1 rejects generate_audio false; Kling 3 prices audio on at $0.21 a second against $0.14 off; recast and motion transfer keep source sound.

5 min readSume
All posts

Not every Sume video model lets you turn the soundtrack off. gemini-omni-flash-1.1 always generates synced audio and the API rejects generate_audio: false. Kling 3 has separate prices with audio on and off. The person-swap model h3-max-recast keeps the source video's sound, and motion transfer takes no generate_audio field either. If your Sora-era pipeline expected a silent clip to lay your own music over, check this before you pick a backend.

What the docs say about the flag

On /v1/videos, generate_audio is a boolean that tells the model to generate audio or not, and its default is the audio capability of the model. Each model in GET /v1/videos/models reports generate_audio, which shows whether it can make an audio track at all. Read that field rather than assuming.

Audio behavior by Sume video model, from the docs and pricing code (read 2026-10-07)
ModelAudio behaviorPrice effect
gemini-omni-flash-1.1Native synced audio always on; false rejectedNone; one price per output second
kling-3Audio on or off$0.14 per second off, $0.21 on (list $0.112 / $0.168 x 1.25)
minimax-h3-maxNative stereo audioPer-second price by resolution
higgsfield-genjutsuNo generate_audio fieldSee the catalog
h3-max-recastNo generate_audio field; keeps sound and cutsPriced by source length

Doing the math on Kling

The pricing code lists Kling 3 Pro at $0.112 per second with audio off and $0.168 with audio on. At Sume's 1.25 multiple that is $0.14 and $0.21. For a 10-second clip, silent costs $1.40 and with sound $2.10, so choosing the silent option saves $0.70 per clip, or $70 across 100 clips. Verify the live figures through the catalog before you rely on them; the code values are list prices from a past date.

If you need silence and the model will not give it

Two options that stay inside Sume. First, choose a model whose generate_audio is controllable and send false. Second, keep the model and drop the audio afterwards: the video trim API has an audio field with keep (default) or drop, and a trim costs $0.02 per job according to its docs. That is a post-step on a clip hosted on media.sume.com, so import the file first if it is not already there.

The second route costs the audio you paid for. On Omni, where audio is not optional, it is the only route.

Dialogue changes the check

If your Sora clips carried dialogue, you want the opposite flag. The dialogue post covers testing short lines on Omni. Whichever way you go, write the audio intent into the prompt and the field together, since the docs treat the flag as the control and not the prompt alone.

Checking before you commit

Run one test per model with the flag you want, read the poll response for usage.cost, and download the file. Then use video inspect to confirm whether an audio stream is present. A catalog flag tells you what a model can do; the inspected file tells you what you got. That two-minute check is cheaper than discovering after a batch that every clip carries a track you did not want.

Sources

Related posts

More in Models

All Models posts

Written by Sume