Sume default queue capacity: max(3, 5 x concurrency), checked per plan
Sume's default queue capacity is max(3, concurrency x 5). Checking the rule against Free, Pro, Startup and Scale, plus the org floor and the wave hint.

Sume's docs give one formula for default queue capacity: max(3, concurrency_limit x 5). Accepted job capacity is that queue plus the processing seats. Checking the formula against the plan table shows it reproduces every default: Free 5, Pro 20, Startup 40, Scale 100, Enterprise 100.
You rarely need the formula, because each submit response carries generation_limits with the effective numbers. But knowing it explains why the queue grows five jobs per seat, and why the wave hint is not a limit.
Formula against the plan table
Concurrency values come from the docs table. The formula column is my calculation from the documented rule; the queue column is the documented value. They agree on every row.
| Plan | Concurrency | max(3, 5 x concurrency) | Documented queue | Accepted jobs |
|---|---|---|---|---|
| Free | 1 | 5 | 5 | 6 |
| Pro | 4 | 20 | 20 | 24 |
| Startup | 8 | 40 | 40 | 48 |
| Scale | 20 | 100 | 100 | 120 |
| Enterprise | 20 | 100 | 100 | 120 |
The floor and the overrides
The max(3, ...) floor only matters if concurrency could be below one, which no plan does; it is a guard. Two other facts do matter. Org workspaces have a floor of 10 on concurrency, and Enterprise defaults to 20 with admin overrides for contract limits. When an override applies, limit_source reads admin_override, and plan_concurrency_limit is not the number to use for sizing.
Prepaid top-ups do not change concurrency, which is plan-only.
The wave hint is not a width
wave_size_hint is max(1, floor(queue_capacity_remaining x 0.75)). On an empty Pro workspace that is floor(24 x 0.75) = 18, which matches the example in the docs. It is a hint for how many jobs to submit in one wave, not a limit and not a processing width.
For in-flight work use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), limited to queue_capacity_remaining. The docs' worked case: 30 processing and 10 queued under a limit of 100 leaves a budget of 60.
def wave_hint(queue_capacity_remaining):
return max(1, int(queue_capacity_remaining * 0.75))
def inflight_budget(limits):
free = max(0, limits["concurrency_limit"] - limits["active_generation_jobs"] - limits["queued_generation_jobs"])
return min(free, limits["queue_capacity_remaining"])
print(wave_hint(24))
print(inflight_budget({"concurrency_limit": 100, "active_generation_jobs": 30, "queued_generation_jobs": 10, "queue_capacity_remaining": 560}))
Takeaway
Read the live fields on every response, and use the formula only to understand them. If the counts are not present on a response, refresh before you choose a width. Remember the counts are a snapshot that other clients and workers can change a moment later.
Using the live fields
The reliable pattern is simple. Submit one request, read generation_limits from the response, compute your budget, and size the next wave from it. Do it again after each wave. If limit_source says admin_override, trust the effective numbers and ignore the plan row.
This matters most for teams with several services sharing one workspace, where each service sees only its own submits but the counts include everyone.
Keep a small table of the last numbers you saw in your logs. When a 429 queue_full arrives, the logged queue_capacity_remaining of the previous response tells you whether you overshot or whether a neighbour service filled the queue. Without the history, both look identical, and teams tend to blame the rate limit for what is really a capacity problem.
Sources
Related posts
More in Developers
- Deno: time out a Sume video submit, then retry with the same key
A fetch that times out may still have created a job. A 20-line Deno submit retries only 408, 429, 5xx and timeouts, always with one Idempotency-Key.
- Idempotency-Key from an order id and version, never a fresh uuid
A fresh uuid per request makes Idempotency-Key do nothing. Derive it from the order id plus a version you bump only to re-run. Scope: one Format, 255 chars.
- How do I narrate a DIY tutorial step by step with a TTS API?
Narrate an 8-step DIY tutorial with one TTS job per step: 1,570 characters, $0.10 on Sume. Why per-step jobs make a fixed step a 1-cent redo.
- Do I pay for a failed AI avatar video job? Refunds on Sume
Sume reserves the avatar video price at submit, captures it on completion, and releases or refunds it where a job fails. What it means for retries.
Written by Sume