Gemini Omni Flash provisioned throughput vs Sume plan concurrency

The Gemini API lists provisioned throughput as unsupported for Omni Flash; Google Cloud says it is rolling out. Sume's capacity is a plan concurrency limit.

5 min readSume
All posts

The Gemini API page for Omni Flash lists provisioned throughput under unsupported features, while Google Cloud's June 30 post said it was rolling out soon on its platform (read 2026-10-02). So reserved capacity for Omni is not something you can count on in the Gemini API today. Sume offers no reserved-capacity product either: capacity is the processing concurrency of your plan.

What Google says

Two pages, two scopes. The Gemini API docs name provisioned throughput among unsupported features for Omni Flash (Gemini API docs, read 2026-10-02). Google Cloud's announcement says provisioned throughput is available for Nano Banana 2 Lite and rolling out for Gemini Omni Flash soon (Google Cloud blog, read 2026-10-02).

Provisioned throughput statements (read 2026-10-02)
SourceModelStatement
Gemini API docsOmni FlashListed as unsupported
Google Cloud blogNano Banana 2 LiteAvailable starting now
Google Cloud blogGemini Omni FlashRolling out soon

What Sume offers

Sume admits paid generation jobs queue-first. Concurrency is a dispatch limit, not a submit limit, so a job can wait as queued until a slot opens, and a full queue returns 429 queue_full (Generation admission).

Processing concurrency is plan-only, and prepaid top-ups do not raise it. Free is 1, Pro 4, Startup 8, Scale 20 and Enterprise 20, with admin overrides possible for contract limits. Read the effective value from generation_limits.concurrency_limit or the dashboard Concurrency tab.

How to plan around it

  • Treat Omni throughput as shared, not reserved, on both sides until Google says otherwise.
  • Spread a launch over the queue: submit everything with idempotency keys and let the window drain.
  • Use a webhook per job so a slow queue does not need your poller running.
  • If you need higher concurrency, that is a plan or Enterprise conversation, not a setting a top-up changes.

The honest comparison

Provisioned throughput on Google's platform is a purchase of dedicated capacity. Sume does not sell a per-model reservation. If your risk is an hourly peak, test your queue depth and plan limit before the peak, since a deeper queue means longer waits rather than rejected jobs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume