gpt-image-2.5 rate tiers (5 to 250 IPM) vs Sume plan queue

OpenAI limits gpt-image-2.5 Flare from 5 images per minute at Tier 1 to 250 at Tier 5. Sume limits processing concurrency by plan and queues the rest.

5 min readSume
All posts

OpenAI's Flare page lists rate limits from 5 images per minute (Tier 1) to 250 (Tier 5), with token limits alongside. Sume has no per-minute image cap in the docs I read; it limits how many jobs process at once by plan, from 1 on Free to 20 on Scale, and holds the rest in a queue.

What are OpenAI's five tiers?

From the gpt-image-2.5 Flare model page, read 2026-10-02. TPM is tokens per minute and IPM is images per minute.

gpt-image-2.5 Flare rate limits (OpenAI, read 2026-10-02)
TierTokens per minuteImages per minute
Tier 1100K5
Tier 2250K20
Tier 3800K50
Tier 43M150
Tier 58M250

How does Sume limit throughput?

Sume separates four controls. Generation concurrency covers jobs in processing. Queue capacity covers accepted but not started jobs. Submit rate limits cover request volume. Balance covers reservation. Concurrency is plan-only: top-ups do not raise it, and queue capacity defaults to max(3, concurrency_limit × 5).

Sume plan limits (Generation admission docs, read 2026-10-02)
PlanProcessing concurrencyQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Can I compare the two directly?

Not cleanly, because they measure different things. An images-per-minute cap depends on how long each render takes. A concurrency cap depends on how many run at once. If a render takes about a minute, 20 concurrent slots would be roughly 20 images a minute; that is arithmetic, not a Sume promise, and OpenAI's page itself says complex prompts can take up to two minutes.

What happens when I exceed the limit?

On Sume, valid jobs beyond concurrency are accepted as queued while capacity remains. A full queue returns 429 queue_full, request volume returns 429 rate_limited, and a short balance returns 402 insufficient_credits. Retry with backoff and an idempotency key. For image calls, POST /v1/images waits up to 30 seconds and then returns 202 with a job, which suits queued work.

The effective number is generation_limits.concurrency_limit; org workspaces have a floor of 10 and Enterprise uses admin overrides, so prefer the field to the table.

Which should I plan around?

If you are a Tier 1 OpenAI account bursting a hundred images, you wait on a 5 IPM cap. On Sume you submit them all, up to accepted capacity, and read results by webhook or poll. If you need more than the accepted capacity at once, you need a higher plan, not a top-up. Neither route publishes a latency guarantee, so measure your own.

See the Tier 1 comparison for the narrow case.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume