OpenAI's three tiers vs Sume plan concurrency of 1, 4, 8 and 20
OpenAI cut API usage tiers from five to three on Oct 6. Sume sets concurrency by plan, not spend. The plan numbers and the queue math, side by side.

OpenAI's API changelog says that on Oct 6 it simplified API usage tiers from five to three, named Build, Launch and Grow, with automatic upgrades when cumulative credit purchases reach a threshold. If you moved a media workload to Sume, the question changes: Sume does not tier by spend. Generation concurrency is set by plan, and prepaid top-ups do not raise it.
This matters when you size a batch. On OpenAI the ceiling you plan around follows how much you have bought. On Sume the ceiling follows the plan on the workspace, and the dashboard Concurrency tab and the generation_limits object in each submit response show the effective number.
The four plan numbers
The table is copied from the Generation admission page. The default queue capacity is max(3, concurrency_limit x 5), and accepted job capacity is concurrency plus queue. Enterprise defaults match Scale and use admin overrides above that.
Org workspaces have a floor of 10 processing slots, per the same page. Treat the static table as a default and read generation_limits.concurrency_limit for the real value, because an admin override changes it and limit_source then says admin_override.
| Plan | Processing | Queue (default) | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
What changes in your client
A queued job is a normal accepted state, not an error. With a Pro workspace you can submit 24 valid paid jobs at once: 4 move to processing, 20 sit in queued, and job 25 fails with 429 queue_full. Nothing is billed for the failed admission, and the reservation is released.
So the client rule is to size waves from generation_limits, not from a tier name. queue_capacity_remaining is the number of accepted slots still open, and wave_size_hint is a submission hint that includes queue slots. Never display the hint as concurrency.
- Read
generation_limitsfrom each submit response and stop adding work whenqueue_capacity_remainingis low. - On
queue_full, wait for a terminal job or cancel queued jobs you no longer need, then retry with the sameIdempotency-Key. - Do not expect a top-up to lift the processing cap. The docs say concurrency is plan-only.
Where the two models differ
OpenAI's change is about how an account graduates between tiers as it buys credits. Sume's balance and concurrency are separate controls: an 402 insufficient_credits means the estimated cost cannot be reserved, while 429 queue_full means the accepted-job capacity is used up. A client that treats both as one 'limit' error will retry the wrong way.
If your job mix is mostly long video, plan around the queue, not the processing slots. A Free workspace holds 6 jobs total, so a batch of 8 gets 6 accepted and 2 rejected. Read the Free plan example for the arithmetic with real clips.
A worked batch
Take 100 video clips on a Pro workspace. The accepted capacity is 24, so submit the first 24, keep them polled, and add one new job each time one reaches a terminal state. Four run at once, so the 100 clips take roughly 25 sequential slot-turns of the longest job, not one. If you submit all 100 at once, jobs 25 to 100 fail with 429 queue_full and you have to resubmit them with the same keys.
On a Scale workspace, 100 clips fit inside the 120 accepted slots, so they can all be accepted in one pass while only 20 run at a time. That is the practical difference between the plans: the queue absorbs the burst, and the processing number sets throughput. OpenAI's tier change affects how quickly an account earns higher ceilings; on Sume you change the plan, and the new numbers appear in generation_limits on the next submit.
Sources
Related posts
More in Developers
- OpenRouter video client on Sume: cancelled is terminal, callback_url
Moving an OpenRouter video client to Sume: add cancelled to your terminal statuses, send callback_url on each request, and keep unknown statuses non-terminal.
- Org workspace concurrency floor of 10: the queue and wave that follow
Sume gives org workspaces a processing floor of 10. With the default queue formula that is 50 queued, 60 accepted, and a wave hint of 45. Read your own fields.
- Pick a Sume video model in code: audio refs, 1080p, 20 seconds
Filter GET /v1/video-router/models on reference_audios, resolutions and duration_seconds, then submit the survivor with reference_audio_urls (up to five).
- Pick the Sume video model in code: seconds, ratio, edit or swap
A 12-line Python function turns duration, aspect ratio, edit and swap needs into the models that can run the job. A 4 s 1:1 clip has none; a 12 s one has two.
Written by Sume