Omni Flash audio is always on, so there is no silent discount on Sume

Gemini Omni Flash 1.1 on Sume always renders native synced audio and rejects generate_audio false. The rate rows have no audio switch to save money.

4 min readSume
All posts

Short answer

On Sume, gemini-omni-flash-1.1 always generates native synced audio. The Video Router docs say the API rejects generate_audio: false, and the model has no bitrate_mode and no reference_audio_urls. Because audio is always on, the price rows are per output second and per resolution only: there is no cheaper silent tier for this model (Video Router docs).

What the rows look like

Sume bills the provider list times 1.25 for each output second as a function of resolution.

Omni Flash 1.1 on Sume, per output second (list x 1.25; read 2026-10-05)
ResolutionPer second10 seconds
360p$0.0375$0.375
720p$0.125$1.25
1080p$0.1875$1.875
4K$0.375$3.75

If your final must be silent

Models differ. The model list in GET /v1/videos/models includes a generate_audio field that shows whether a model can generate an audio track; for example the seedance-2 entry in the docs example has generate_audio: true and accepts an audio reference. Read the field for each model before you plan a silent deliverable, because the Omni field is always on.

The Google launch post for Gemini Omni 1.1 Flash (read 2026-10-05) does not state an audio price on the page I read, so the Sume rows are the numbers to budget with.

Using the audio on purpose

  • Write the sound into the prompt; the audio is generated with the picture, not added later.
  • Use the Sume Timeline docs if you need to replace the soundtrack with your own audio spine in the final render.
  • Budget the whole clip at one price per second; audio never adds a line item on this model.

A worked budget with audio included

Five 8-second 720p clips are 5 x 8 x $0.125 = $5.00 with audio. There is no line item to remove. If you need to compare against a pipeline that generates silent video and adds voice separately, price both honestly: the silent video from another model plus a text-to-speech job plus a timeline render, against the Omni clip where audio is bundled. Sume lists those building blocks, and the stored comparisons elsewhere on this blog work through them.

The core point stands: with Omni there is one number to budget, and it includes sound.

For planning, this simplifies one thing and complicates another. It simplifies the budget, because there is no audio toggle that changes the bill. It complicates workflows that need a silent clip, since you would either choose another model or mute the track downstream. If the rest of your timeline carries its own sound, check how the Omni audio sits against it before you commit a large batch.

A quick test with a short clip at 360p costs a few cents and shows you the audio the model produces, which is the only reliable way to judge whether you want it.

The docs state the rule directly: generate_audio: false is rejected for this model, so a request that tries to turn sound off fails validation instead of being billed at a different rate. That is useful behavior, because it means you cannot accidentally pay for audio you thought you removed, or expect a discount that does not exist. If a client library sets that flag by default for other models, check it before you point the same code at Omni.

Sources

Related posts

More in Models

All Models posts

Written by Sume