Python: check Omni Flash 1.1 limits against /v1/videos/models first

Google's Omni runs a sync call and extends in 10-second steps up to 40 seconds. Sume's gemini-omni-flash-1.1 takes 3 to 10 seconds per job. A preflight check.

4 min readSume
All posts

Fetch GET /v1/videos/models, find gemini-omni-flash-1.1, and compare your duration, resolution and aspect ratio with its supported_* lists before you submit. A 12-second request fails that check, since Sume lists 3 to 10 seconds for the model. Google's Omni page describes extension in steps of up to 10 seconds with a 40-second cap on the total, which is a different mechanism from a single longer request.

Google's page and Sume's catalog

Google's Omni docs say generation is synchronous and unary by default, with a response under 4 MB returned as base64 and larger ones fetched by URI after the file turns ACTIVE. Sume wraps the same family in an async job, so there is no 4 MB branch to handle: you poll or take a webhook, then read the result.

Omni on Google's API vs gemini-omni-flash-1.1 on Sume (Google Omni page and Sume docs, read 2026-10-07)
ItemGoogle Omni pageSume catalog entry
Resolutions360p, 720p, 1080p (upscaled), 4K (upscaled)360p, 720p, 1080p, 4K
Aspect ratios16:9 or 9:1616:9 or 9:16
LengthExtension up to 10 s per operation, 40 s total3 to 10 seconds per job
Call styleSynchronous unary by defaultAsync job, 202 with an id
Idgemini-omni-1.1-flashgemini-omni-flash-1.1

The preflight

The script reads the catalog once and checks each field you pass. It uses the standard library and asyncio.to_thread, so the fetch does not block an event loop, and it ends with asyncio.run(main()). A None result means every field is in range; otherwise you get the first mismatch and the allowed list.

import asyncio
import json
import os
import urllib.request

BASE = os.environ.get("SUME_BASE", "https://api.sume.com")

def fetch_models():
    req = urllib.request.Request(BASE + "/v1/videos/models", headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=10) as res:
        return json.load(res)["data"]

def check(models, model_id, **want):
    m = next((m for m in models if m["id"] == model_id), None)
    if m is None:
        return f"{model_id} is not in GET /v1/videos/models"
    for field, value in want.items():
        allowed = m["supported_" + field + "s"]
        if value not in allowed:
            return f"{field}={value} not supported by {model_id}: {allowed}"

async def main():
    models = await asyncio.to_thread(fetch_models)
    print(check(models, "gemini-omni-flash-1.1", duration=12, resolution="4K"))
    print(check(models, "gemini-omni-flash-1.1", duration=10, aspect_ratio="9:16"))

asyncio.run(main())

Limits

The check only reads the catalog; it does not know price or whether the provider is up. The catalog can change, so read it at run time rather than copying numbers into code. Sume does not offer Google's extension flow as a single request. To reach more than 10 seconds, plan several jobs, as the linked chaining post explains.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume