Silent clips after Sora: which Sume models take generate_audio false
Gemini Omni Flash 1.1 rejects generate_audio false; Kling 3 prices audio on at $0.21 a second against $0.14 off; recast and motion transfer keep source sound.

Not every Sume video model lets you turn the soundtrack off. gemini-omni-flash-1.1 always generates synced audio and the API rejects generate_audio: false. Kling 3 has separate prices with audio on and off. The person-swap model h3-max-recast keeps the source video's sound, and motion transfer takes no generate_audio field either. If your Sora-era pipeline expected a silent clip to lay your own music over, check this before you pick a backend.
What the docs say about the flag
On /v1/videos, generate_audio is a boolean that tells the model to generate audio or not, and its default is the audio capability of the model. Each model in GET /v1/videos/models reports generate_audio, which shows whether it can make an audio track at all. Read that field rather than assuming.
| Model | Audio behavior | Price effect |
|---|---|---|
| gemini-omni-flash-1.1 | Native synced audio always on; false rejected | None; one price per output second |
| kling-3 | Audio on or off | $0.14 per second off, $0.21 on (list $0.112 / $0.168 x 1.25) |
| minimax-h3-max | Native stereo audio | Per-second price by resolution |
| higgsfield-genjutsu | No generate_audio field | See the catalog |
| h3-max-recast | No generate_audio field; keeps sound and cuts | Priced by source length |
Doing the math on Kling
The pricing code lists Kling 3 Pro at $0.112 per second with audio off and $0.168 with audio on. At Sume's 1.25 multiple that is $0.14 and $0.21. For a 10-second clip, silent costs $1.40 and with sound $2.10, so choosing the silent option saves $0.70 per clip, or $70 across 100 clips. Verify the live figures through the catalog before you rely on them; the code values are list prices from a past date.
If you need silence and the model will not give it
Two options that stay inside Sume. First, choose a model whose generate_audio is controllable and send false. Second, keep the model and drop the audio afterwards: the video trim API has an audio field with keep (default) or drop, and a trim costs $0.02 per job according to its docs. That is a post-step on a clip hosted on media.sume.com, so import the file first if it is not already there.
The second route costs the audio you paid for. On Omni, where audio is not optional, it is the only route.
Dialogue changes the check
If your Sora clips carried dialogue, you want the opposite flag. The dialogue post covers testing short lines on Omni. Whichever way you go, write the audio intent into the prompt and the field together, since the docs treat the flag as the control and not the prompt alone.
Checking before you commit
Run one test per model with the flag you want, read the poll response for usage.cost, and download the file. Then use video inspect to confirm whether an audio stream is present. A catalog flag tells you what a model can do; the inspected file tells you what you got. That two-minute check is cheaper than discovering after a batch that every clip carries a track you did not want.
Sources
Related posts
More in Models
- A 60-second Seedance 2.5 piece: two 30 s jobs and one last frame
Together lists multi-round extension for Seedance 2.5; Sume documents none. Build 60 s as two 30 s jobs and a last-frame handoff: $34.68 at 720p, $7.50 on Wan.
- Square 1:1 ad video: which Sume models accept it, and 4:5
Seedance, Kling, Wan and MiniMax accept 1:1 on Sume; Gemini Omni takes only 16:9 and 9:16. No video model lists 4:5, so Feed needs a crop. Table and steps.
- MiniMax H3 Max on Sume: video, lip-sync and recast, which id to call
Three Sume ids carry the MiniMax H3 Max name: minimax-h3-max for video, a lip-sync route and h3-max-recast. What each takes, its window and its rate.
- Two reference images, one clip: Omni's cat-and-yarn pattern on Sume
Google's docs show two images, a cat and yarn, producing one video. Here is the Sume request with IMAGE_REF tokens and what changes with a brand product.
Written by Sume