gpt-image-2.5 rate tiers (5 to 250 IPM) vs Sume plan queue
OpenAI limits gpt-image-2.5 Flare from 5 images per minute at Tier 1 to 250 at Tier 5. Sume limits processing concurrency by plan and queues the rest.

OpenAI's Flare page lists rate limits from 5 images per minute (Tier 1) to 250 (Tier 5), with token limits alongside. Sume has no per-minute image cap in the docs I read; it limits how many jobs process at once by plan, from 1 on Free to 20 on Scale, and holds the rest in a queue.
What are OpenAI's five tiers?
From the gpt-image-2.5 Flare model page, read 2026-10-02. TPM is tokens per minute and IPM is images per minute.
| Tier | Tokens per minute | Images per minute |
|---|---|---|
| Tier 1 | 100K | 5 |
| Tier 2 | 250K | 20 |
| Tier 3 | 800K | 50 |
| Tier 4 | 3M | 150 |
| Tier 5 | 8M | 250 |
How does Sume limit throughput?
Sume separates four controls. Generation concurrency covers jobs in processing. Queue capacity covers accepted but not started jobs. Submit rate limits cover request volume. Balance covers reservation. Concurrency is plan-only: top-ups do not raise it, and queue capacity defaults to max(3, concurrency_limit × 5).
| Plan | Processing concurrency | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
Can I compare the two directly?
Not cleanly, because they measure different things. An images-per-minute cap depends on how long each render takes. A concurrency cap depends on how many run at once. If a render takes about a minute, 20 concurrent slots would be roughly 20 images a minute; that is arithmetic, not a Sume promise, and OpenAI's page itself says complex prompts can take up to two minutes.
What happens when I exceed the limit?
On Sume, valid jobs beyond concurrency are accepted as queued while capacity remains. A full queue returns 429 queue_full, request volume returns 429 rate_limited, and a short balance returns 402 insufficient_credits. Retry with backoff and an idempotency key. For image calls, POST /v1/images waits up to 30 seconds and then returns 202 with a job, which suits queued work.
The effective number is generation_limits.concurrency_limit; org workspaces have a floor of 10 and Enterprise uses admin overrides, so prefer the field to the table.
Which should I plan around?
If you are a Tier 1 OpenAI account bursting a hundred images, you wait on a 5 IPM cap. On Sume you submit them all, up to accepted capacity, and read results by webhook or poll. If you need more than the accepted capacity at once, you need a higher plan, not a top-up. Neither route publishes a latency guarantee, so measure your own.
See the Tier 1 comparison for the narrow case.
Sources
Related posts
More in Comparisons
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
- Hedra 402 INSUFFICIENT_BALANCE vs Sume insufficient_credits
Hedra returns 402 INSUFFICIENT_BALANCE until you add funds; Sume returns 402 insufficient_credits. How each wallet check behaves and how to preflight a job.
- Hedra Avatar needs start frame and audio; Sume takes a handle and scri
Hedra Avatar generates from a start frame plus an audio track, up to 10 minutes. Sume's talking-video takes an avatar handle and a script, up to 60 seconds.
- Hedra's developer platform: API, SDK, CLI and MCP vs Sume
Hedra opened its models through an API, SDKs, a CLI and MCP on August 4, 2026. A map of what each surface covers and what Sume offers for avatar work.
Written by Sume