Silent 11-second b-roll: Omni rejects generate_audio false
Gemini Omni Flash 1.1 on Sume always has audio and rejects generate_audio false, and stops at 10 s anyway. For 11 s of b-roll, Wan 3.0 at 720p is $1.375.

Gemini Omni Flash 1.1 cannot make silent b-roll on Sume: native audio is always on and the API rejects generate_audio: false. It also stops at 10 s, so an 11-second clip is out on both counts. For 11 s at 720p, Wan 3.0 costs 11 s × $0.125 = $1.375 and Seedance 2.5 11 s × $0.5778 = $6.3558.
Prices on this page are Sume list prices as of 2026-10-08: the provider list price times 1.25, computed from Sume's pricing package and billed per output second, linear in duration inside each model's valid range. The model ids, ranges and input types come from the Video generation docs and the Video Router docs; the per-second numbers can be cross-checked in the public catalog.
How to check a model's audio switch
Every model in GET /v1/videos/models reports generate_audio, which shows whether it can generate an audio track, and the request field of the same name is described as telling the model to generate audio or not, with the default being the model's own capability. Read the flag for the model you pin instead of assuming; this page does not claim a silent mode for any model other than what the docs state.
If you cannot get a silent render, the fallback is to remove the track afterwards. The catalog lists audio detach at $0.01 per job.
| Model | 11 s allowed | 11 s price | Audio |
|---|---|---|---|
| Gemini Omni Flash 1.1 | No (3 to 10 s) | n/a | always on |
| Wan 3.0 | Yes (2 to 30 s) | $1.375 | check generate_audio |
| Seedance 2.5 | Yes (4 to 30 s) | $6.3558 | check generate_audio |
| Seedance 2.0 | Yes (4 to 15 s) | $4.158 | generate_audio: true in the docs example |
| MiniMax H3 Max | Yes (5 to 15 s) | $1.10 | native stereo audio |
Cost of the fix
Detaching audio from an 11 s Wan 3.0 clip adds $0.01, so a silent clip is $1.385 if you have to strip it. For most b-roll the narration or music goes over it, so the point is mostly that Omni's own sound will be present under it.
Scaling this plan up or down
As a yardstick, the plan above is built on Wan 3.0 at 720p, $0.125 per second. Each extra 5 seconds adds $0.625, a further $10 of budget buys 80 more whole seconds, and the largest single job the model accepts (30 s) holds $3.75 at submit. The shortest one (2 s) holds $0.25.
Those three numbers are enough to rescale the plan without a new table. If the plan doubles, double the totals; if a clip is shortened, subtract the seconds multiplied by the rate; and if the tier changes, swap the rate for the one in the tables above.
- Per second: $0.125
- Per 5 s: $0.625
- Per 10 s: $1.25
- Per 30 s or the model maximum (30 s): $3.75
Check the model's fields first
Before a run, ask GET /v1/videos/models for the model and compare three fields with your plan: supported_durations (every clip length must be listed), supported_resolutions (the tier must be listed, since MiniMax H3 has 768p and no 1080p, and Omni has 360p and 4K) and supported_aspect_ratios. A request outside those lists fails at validation, so the failure costs nothing, but a plan built on a wrong assumption costs a rewrite.
Submitting
Submit through POST /v1/videos with model, prompt, duration in whole seconds and resolution; read GET /v1/videos/models first, because supported_durations and supported_resolutions differ per model. Sume reserves the price at submit and usage.cost on the poll response is the billable amount, so the figures here are what leaves the balance.
Sources
Related posts
More in Developers
- Silent b-roll Short: Timeline audio.mode silence, 179 s for $0.30
Render a silent vertical Short with Timeline 1.0 audio.mode silence: 179 s bills 3 minutes ($0.30), and four refusal codes say which field to drop.
- 60 video jobs in Node with a promise pool: which plans need waves
A Node promise pool for Sume video jobs, plus the fit table: 60 jobs against accepted capacity on Free, Pro, Startup and Scale, and how many waves each needs.
- Smallest square GPT Image 2.5 accepts on Sume: 816x816
A small icon request fails at 800x800 on GPT Image 2.5: under 655,360 pixels. 810x810 fails the 16 rule. 816x816 is the smallest legal square.
- sonic-latest or sonic-3-6: which Sume TTS id for repeatable narration
Pin sonic-3.6 when a series must sound the same; sonic-latest is an alias that tracks the current stable Sonic. Avoid sonic-preview with pro voice clones.
Written by Sume