Gemini tiers unlock by spend and days. Sume limits follow the plan

Gemini upgrades a project after spend and elapsed days. Sume sets request and concurrency limits by plan, and a prepaid top-up does not raise concurrency.

4 min readSume
All posts

Gemini raises a project's rate limits as it meets spend and time thresholds, while Sume ties limits to the subscription plan. Google's page lists Tier 2 as $100+ spending plus 3 days elapsed and Tier 3 as $1,000+ plus 30 days. Sume's docs say processing concurrency is plan-only and that prepaid top-ups do not raise it.

Google facts are from its rate limits page (listed under Sources) and Sume facts from Generation admission and Authentication, read 2026-09-30.

How does each side decide your limits?

Google says tiers upgrade automatically once a project meets the qualification criteria. Sume reads the plan of the workspace the key belongs to.

Concurrency by plan from Generation admission and Authentication, read 2026-09-30; always prefer the effective field
PlanProcessing concurrencyQueue capacityWrites per minute
Free15120
Pro420300
Startup840600
Scale201001200

Will topping up credits lift my concurrency?

No. The docs say generation concurrency is plan-only and prepaid top-ups do not raise the processing limit. Admin overrides can raise the effective concurrency_limit (limit_source: admin_override), and organization workspaces have a floor of 10. The dashboard Concurrency tab is the source of truth, exposed as generation_limits.concurrency_limit.

What does the queue do when concurrency is full?

It accepts valid jobs as queued while queue capacity remains and moves them to processing as slots open. Only when accepted capacity (concurrency plus queue) is full does a submit fail with 429 queue_full. Enterprise request limits are contract-specific; until provisioned, an Enterprise key resolves to the Scale row.

What are the limits of this comparison?

The two systems measure different things: Gemini limits are requests and tokens per project, Sume's headline limit here is concurrent generation jobs and requests per key. The static table is the shipped default; read your own generation_limits before planning a batch.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume