Free, Pro, Startup, Scale: processing seats, queue slots, full hold
Sume's concurrency by plan, queue capacity max(3, 5 x concurrency), accepted job capacity, and the balance reserved if every slot holds a 10 s clip.

Sume admits paid generation jobs by plan: Free runs 1 at a time with 5 queue slots, Pro 4 with 20, Startup 8 with 40, and Scale and Enterprise 20 with 100. Queue capacity defaults to max(3, concurrency_limit x 5), and accepted capacity is concurrency plus queue, so a Pro workspace holds 24 paid jobs at once before 429 queue_full.
Concurrency limits dispatch, not submission. A Pro workspace can submit 24 valid jobs in one burst; four run and twenty wait as queued. Generation admission is the source for every number here, and the API's generation_limits field is the live truth if your workspace has an override.
Capacity and worst-case hold
The last column is the reserve if every accepted slot holds a 10-second Wan 3.0 clip at 720p ($1.25 each). Sume reserves each estimate at submit.
| Plan | Processing | Queue slots | Accepted | Hold if all full |
|---|---|---|---|---|
| Free | 1 | 5 | 6 | $7.50 |
| Pro | 4 | 20 | 24 | $30.00 |
| Startup | 8 | 40 | 48 | $60.00 |
| Scale | 20 | 100 | 120 | $150.00 |
| Enterprise | 20 | 100 | 120 | $150.00 |
Reading the live numbers
Submit responses carry a generation_limits snapshot. The docs size new in-flight work from it:
def headroom(g):
"""New in-flight jobs to add, per the generation-admission docs."""
free_seats = g["concurrency_limit"] - g["active_generation_jobs"] - g["queued_generation_jobs"]
return min(max(0, free_seats), g["queue_capacity_remaining"])
snapshot = {"concurrency_limit": 4, "active_generation_jobs": 1,
"queued_generation_jobs": 0, "queue_capacity_remaining": 23}
print(headroom(snapshot)) # 3Two traps
wave_size_hintis a submission hint that includes queue slots. It is not your concurrency, so never size a worker pool from it.- Concurrency is plan-only: prepaid top-ups do not raise it. Org workspaces have a floor of 10 and Enterprise uses admin overrides.
How to use the table
The hold column is a ceiling, not a typical state. It answers one question: what balance would a worst-case burst reserve? A workspace that has accepted its full capacity of paid jobs has that much set aside until jobs complete, fail or are canceled, so a balance below the figure means some submits in a burst would be refused with 402 insufficient_credits before the queue is even full.
Planning a burst
Take your accepted capacity, multiply by the per-job estimate, and compare with the balance. If the balance is the binding limit, you will see 402s first; if the queue is, you will see 429 queue_full. Either way the right pattern is the same: submit in waves, store each job id, and let finished jobs free room for the next wave. The Python helper above computes how many jobs a wave can add from a live snapshot.
Sources
Related posts
More in Developers
- Typed Sume video client from the OpenAPI JSON, after Sora
Sume publishes OpenAPI 3.0.3 at api.sume.com/reference/json. List the four video operations, then use the SDK's generated calls instead of hand-typing the wire.
- Generate then cut out: an image-model result into RMBG, in Python
Two Sume calls: generate a product shot, then POST its URL to /v1/rmbg-1.0/remove. Runnable Python, the polling loop, and the $0.1225 total per cutout.
- Go: net/http client for a 30-second Wan 3.0 clip, 30 lines
A 30-line Go program using only the standard library: submit wan-3.0 for 30 seconds, poll the job, write clip.mp4. Reserve and per-second rate included.
- gpt-image-2.5 quality auto reserves max: holds from $0.22 to $0.89
On Sume, gpt-image-2.5 with quality auto and auto size reserves $0.8895 per image, while omitting quality reserves $0.2224. Hold table and the safe request.
Written by Sume