Fields Omni rejects on Sume: generate_audio false, bitrate_mode
Which request fields Sume's gemini-omni-flash-1.1 refuses or lacks: generate_audio false, bitrate_mode, reference_audio_urls, plus edit-mode rules.

Sume's gemini-omni-flash-1.1 rejects generate_audio: false, because native audio is always on, and it has no bitrate_mode and no reference_audio_urls. If your code was written for another video model and sends those fields, remove them for Omni. The authority is GET /v1/video-router/models (or GET /v1/videos/models), which reports the capabilities of each model; Sume's Video Router docs state these limits too (Sume docs: Video Router, read 2026-10-05).
What the docs say about Omni's envelope
Sume routes by the shape of the request. You never choose an endpoint.
| Request shape | You send | Notes |
|---|---|---|
| Text to video | prompt | 3 to 10 s, 360p to 4K, 16:9 or 9:16 |
| Image to video | image_url, optional end_image_url | Same envelope |
| Reference to video | reference_image_urls (up to 10) and/or reference_video_urls (up to 3, each up to 3 s) | Name them <IMAGE_REF_0> and <VIDEO_REF_0> in the prompt |
| Edit | video_url | Resolution optional, default 720p, no aspect_ratio and no duration |
Fields that fail on Omni
Do not guess which fields a model takes. Read capabilities from the models endpoint for the model you pin, because each model has its own envelope.
generate_audio: false: refused. Native audio is always on. If you need silence, remove the audio in a later step, such as the timeline's silence spine or a trim withaudio: drop.bitrate_mode: not offered for this model.reference_audio_urls: not offered. Voice and sound come from the prompt, not from an audio reference.durationandaspect_ratioon an edit: do not send them. The edit keeps the source clip's length and shape.- Durations outside 3 to 10 seconds: out of range. Ask for 10 and chain if you need longer.
Fields that are not there at all
Sume does not give Omni an extend mode. Google's Omni 1.1 Flash supports scene extension in 10-second steps up to 40 seconds (Google, read 2026-10-05), but the Sume catalog entry has the four shapes in the table. To go past 10 seconds on Sume, make several jobs and join them with a timeline, using the last frame of one clip as image_url of the next.
A defensive client strips unknown fields per model before it submits. Keep a small allow-list per model, built from the models endpoint, and log what you removed. A silent drop is how you end up with audio you did not want.
Here is a worked example. A client built for a model that takes a bitrate setting and an audio reference sends both to Omni. The audio reference is not offered, so the request fails instead of running with part of your intent. Remove the field, move the sound direction into the prompt, run the job, and keep a note in the code saying why the field is absent so the next person does not add it back.
Handle the refusal as a 400-class error. Do not retry it: a malformed body fails the same way every time, and a retry only wastes a request against your rate limit. Read the error code, fix the body, and resubmit with a new Idempotency-Key, since the old key is tied to the old body.
Last, test the allow-list. Write a unit test that builds a body for each request shape in the table and checks that none of the removed fields is present. It runs without the network and catches the regression the day someone copies a body from another model.
Sources
Related posts
More in Developers
- Filter the Sume video catalog in Python for 1080p and audio references
Instead of guessing which video model takes what, read GET /v1/videos/models and filter it. This Python script lists models with 1080p and audio_url references.
- Find gpt-image-1 in your repo before Oct 23: a Python scan
OpenAI shuts gpt-image-1 down on 2026-10-23, five weeks before the other GPT Image ids. A Python scan splits them by date and maps each hit to a Sume model id.
- Find video models with 1:1, 9:16 and 16:9: filter the Sume catalog
One Python call to GET /v1/videos/models lists which Sume video models cover square, vertical and widescreen, so one prompt can feed Feed, Reels and in-stream.
- Firebase allows 1-hour HTTP functions; don't hold a Sume job open
Firebase HTTP functions can run up to an hour and Cloud Tasks covers longer work. Even so, submit Sume jobs async and take the result by webhook or status poll.
Written by Sume