Gemini Omni Flash provisioned throughput vs Sume plan concurrency
The Gemini API lists provisioned throughput as unsupported for Omni Flash; Google Cloud says it is rolling out. Sume's capacity is a plan concurrency limit.

The Gemini API page for Omni Flash lists provisioned throughput under unsupported features, while Google Cloud's June 30 post said it was rolling out soon on its platform (read 2026-10-02). So reserved capacity for Omni is not something you can count on in the Gemini API today. Sume offers no reserved-capacity product either: capacity is the processing concurrency of your plan.
What Google says
Two pages, two scopes. The Gemini API docs name provisioned throughput among unsupported features for Omni Flash (Gemini API docs, read 2026-10-02). Google Cloud's announcement says provisioned throughput is available for Nano Banana 2 Lite and rolling out for Gemini Omni Flash soon (Google Cloud blog, read 2026-10-02).
| Source | Model | Statement |
|---|---|---|
| Gemini API docs | Omni Flash | Listed as unsupported |
| Google Cloud blog | Nano Banana 2 Lite | Available starting now |
| Google Cloud blog | Gemini Omni Flash | Rolling out soon |
What Sume offers
Sume admits paid generation jobs queue-first. Concurrency is a dispatch limit, not a submit limit, so a job can wait as queued until a slot opens, and a full queue returns 429 queue_full (Generation admission).
Processing concurrency is plan-only, and prepaid top-ups do not raise it. Free is 1, Pro 4, Startup 8, Scale 20 and Enterprise 20, with admin overrides possible for contract limits. Read the effective value from generation_limits.concurrency_limit or the dashboard Concurrency tab.
How to plan around it
- Treat Omni throughput as shared, not reserved, on both sides until Google says otherwise.
- Spread a launch over the queue: submit everything with idempotency keys and let the window drain.
- Use a webhook per job so a slow queue does not need your poller running.
- If you need higher concurrency, that is a plan or Enterprise conversation, not a setting a top-up changes.
The honest comparison
Provisioned throughput on Google's platform is a purchase of dedicated capacity. Sume does not sell a per-model reservation. If your risk is an hourly peak, test your queue depth and plan limit before the peak, since a deeper queue means longer waits rather than rejected jobs.
Sources
Related posts
More in Developers
- Gemini prefixItems tuple schema: Sume rejects it, use an object
Gemini lists prefixItems for tuple-like arrays. Sume's output_schema allowlist omits it and returns unsupported_keyword; model each slot as a named property.
- Gemini recursive schema with $ref "#": what Sume accepts instead
Gemini's docs show an org-chart schema that recurses with $ref "#". Sume rejects that root reference; recurse through a named $defs entry instead.
- Check GET /v1/balance before a Sume bulk run: failures land per item
A bulk create returns 202 even if the wallet cannot fund every child; unfunded items fail one by one. Compare GET /v1/balance with your spend caps first.
- Smoke-test a new Sume API key with GET /v1/me before revoking the old
GET /v1/me returns the key's id, prefix and scopes. A TypeScript script fails the deploy if the new key lacks formats:write, so you revoke the old key after.
Written by Sume