Sume default queue capacity: max(3, 5 x concurrency), checked per plan

Sume's default queue capacity is max(3, concurrency x 5). Checking the rule against Free, Pro, Startup and Scale, plus the org floor and the wave hint.

4 min readSume
All posts

Sume's docs give one formula for default queue capacity: max(3, concurrency_limit x 5). Accepted job capacity is that queue plus the processing seats. Checking the formula against the plan table shows it reproduces every default: Free 5, Pro 20, Startup 40, Scale 100, Enterprise 100.

You rarely need the formula, because each submit response carries generation_limits with the effective numbers. But knowing it explains why the queue grows five jobs per seat, and why the wave hint is not a limit.

Formula against the plan table

Concurrency values come from the docs table. The formula column is my calculation from the documented rule; the queue column is the documented value. They agree on every row.

Queue capacity by plan, documented value vs the documented formula (read 2026-10-07)
PlanConcurrencymax(3, 5 x concurrency)Documented queueAccepted jobs
Free1556
Pro4202024
Startup8404048
Scale20100100120
Enterprise20100100120

The floor and the overrides

The max(3, ...) floor only matters if concurrency could be below one, which no plan does; it is a guard. Two other facts do matter. Org workspaces have a floor of 10 on concurrency, and Enterprise defaults to 20 with admin overrides for contract limits. When an override applies, limit_source reads admin_override, and plan_concurrency_limit is not the number to use for sizing.

Prepaid top-ups do not change concurrency, which is plan-only.

The wave hint is not a width

wave_size_hint is max(1, floor(queue_capacity_remaining x 0.75)). On an empty Pro workspace that is floor(24 x 0.75) = 18, which matches the example in the docs. It is a hint for how many jobs to submit in one wave, not a limit and not a processing width.

For in-flight work use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), limited to queue_capacity_remaining. The docs' worked case: 30 processing and 10 queued under a limit of 100 leaves a budget of 60.

def wave_hint(queue_capacity_remaining):
    return max(1, int(queue_capacity_remaining * 0.75))


def inflight_budget(limits):
    free = max(0, limits["concurrency_limit"] - limits["active_generation_jobs"] - limits["queued_generation_jobs"])
    return min(free, limits["queue_capacity_remaining"])


print(wave_hint(24))
print(inflight_budget({"concurrency_limit": 100, "active_generation_jobs": 30, "queued_generation_jobs": 10, "queue_capacity_remaining": 560}))

Takeaway

Read the live fields on every response, and use the formula only to understand them. If the counts are not present on a response, refresh before you choose a width. Remember the counts are a snapshot that other clients and workers can change a moment later.

Using the live fields

The reliable pattern is simple. Submit one request, read generation_limits from the response, compute your budget, and size the next wave from it. Do it again after each wave. If limit_source says admin_override, trust the effective numbers and ignore the plan row.

This matters most for teams with several services sharing one workspace, where each service sees only its own submits but the counts include everyone.

Keep a small table of the last numbers you saw in your logs. When a 429 queue_full arrives, the logged queue_capacity_remaining of the previous response tells you whether you overshot or whether a neighbour service filled the queue. Without the history, both look identical, and teams tend to blame the rate limit for what is really a capacity problem.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume