OpenAI's three tiers vs Sume plan concurrency of 1, 4, 8 and 20

OpenAI cut API usage tiers from five to three on Oct 6. Sume sets concurrency by plan, not spend. The plan numbers and the queue math, side by side.

5 min readSume
All posts

OpenAI's API changelog says that on Oct 6 it simplified API usage tiers from five to three, named Build, Launch and Grow, with automatic upgrades when cumulative credit purchases reach a threshold. If you moved a media workload to Sume, the question changes: Sume does not tier by spend. Generation concurrency is set by plan, and prepaid top-ups do not raise it.

This matters when you size a batch. On OpenAI the ceiling you plan around follows how much you have bought. On Sume the ceiling follows the plan on the workspace, and the dashboard Concurrency tab and the generation_limits object in each submit response show the effective number.

The four plan numbers

The table is copied from the Generation admission page. The default queue capacity is max(3, concurrency_limit x 5), and accepted job capacity is concurrency plus queue. Enterprise defaults match Scale and use admin overrides above that.

Org workspaces have a floor of 10 processing slots, per the same page. Treat the static table as a default and read generation_limits.concurrency_limit for the real value, because an admin override changes it and limit_source then says admin_override.

Sume default generation limits by plan, as of 2026-10-09 (docs.sume.com/workflows/generation-admission)
PlanProcessingQueue (default)Accepted jobs
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

What changes in your client

A queued job is a normal accepted state, not an error. With a Pro workspace you can submit 24 valid paid jobs at once: 4 move to processing, 20 sit in queued, and job 25 fails with 429 queue_full. Nothing is billed for the failed admission, and the reservation is released.

So the client rule is to size waves from generation_limits, not from a tier name. queue_capacity_remaining is the number of accepted slots still open, and wave_size_hint is a submission hint that includes queue slots. Never display the hint as concurrency.

  • Read generation_limits from each submit response and stop adding work when queue_capacity_remaining is low.
  • On queue_full, wait for a terminal job or cancel queued jobs you no longer need, then retry with the same Idempotency-Key.
  • Do not expect a top-up to lift the processing cap. The docs say concurrency is plan-only.

Where the two models differ

OpenAI's change is about how an account graduates between tiers as it buys credits. Sume's balance and concurrency are separate controls: an 402 insufficient_credits means the estimated cost cannot be reserved, while 429 queue_full means the accepted-job capacity is used up. A client that treats both as one 'limit' error will retry the wrong way.

If your job mix is mostly long video, plan around the queue, not the processing slots. A Free workspace holds 6 jobs total, so a batch of 8 gets 6 accepted and 2 rejected. Read the Free plan example for the arithmetic with real clips.

A worked batch

Take 100 video clips on a Pro workspace. The accepted capacity is 24, so submit the first 24, keep them polled, and add one new job each time one reaches a terminal state. Four run at once, so the 100 clips take roughly 25 sequential slot-turns of the longest job, not one. If you submit all 100 at once, jobs 25 to 100 fail with 429 queue_full and you have to resubmit them with the same keys.

On a Scale workspace, 100 clips fit inside the 120 accepted slots, so they can all be accepted in one pass while only 20 run at a time. That is the practical difference between the plans: the queue absorbs the burst, and the processing number sets throughput. OpenAI's tier change affects how quickly an account earns higher ceilings; on Sume you change the plan, and the new numbers appear in generation_limits on the next submit.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume