Why concurrency_limit differs from the Sume plan table
In Sume generation_limits, concurrency_limit is the effective cap and limit_source says plan or admin_override. Size work from it, not plan_concurrency_limit.

If the generation_limits block on a Sume submit response shows a concurrency_limit that differs from the plan table in the docs, trust concurrency_limit. It is the effective cap, and limit_source tells you where it came from: plan for the plan default, admin_override when Sume raised it for the workspace.
The field next to it, plan_concurrency_limit, is only the plan default. The Generation admission page says to ignore it for wave sizing when an override applies.
Which fields tell me the real limit?
A submit response includes a snapshot like this when Sume can compute it:
concurrency_limitis the effective maximum number of same-workspace paid generation jobs that can beprocessingat once.limit_sourceisplanoradmin_override.plan_concurrency_limitis the plan default and can differ from the effective value.accepted_generation_jobs_limitisconcurrency_limit + queued_jobs_limit.
{
"generation_limits": {
"plan_id": "pro",
"limit_source": "plan",
"plan_concurrency_limit": 4,
"concurrency_limit": 4,
"queued_jobs_limit": 20,
"accepted_generation_jobs_limit": 24
}
}What are the plan defaults, and why do they not always match?
The docs list static defaults per plan. They are a starting point, and the page itself says to prefer the effective field over the table.
Three things can move your effective number. An admin override raises it (limit_source: admin_override). Org workspaces have a floor of 10. And Enterprise defaults to 20 and uses admin overrides for higher contract limits.
| Plan | Processing concurrency | Queue capacity (default) | Accepted job capacity |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
Do top-ups raise my concurrency?
No. Generation concurrency is plan-only. Prepaid top-ups do not raise the processing concurrency limit, so adding funds will not change concurrency_limit. The dashboard Concurrency tab is the source of truth for the workspace's configured cap, exposed as generation_limits.concurrency_limit.
Queue capacity defaults to max(3, concurrency_limit × 5). With an override that raises concurrency, the default queue grows with it, and queued_jobs_limit in the response shows the figure that applies to you.
Why does my dashboard number not match this post?
Because the post quotes defaults. The Concurrency tab shows your workspace's configured processing cap, and the API exposes the same figure as generation_limits.concurrency_limit. If the two ever disagree, the docs name the API field as the effective one.
Also keep purchased_concurrency_limit-style fields out of your sizing logic. The docs say never to substitute the pre-override fields for the effective one, and top-ups do not change processing concurrency at all.
Finally, remember what the cap controls. It limits how many generation jobs are processing at once, not how many you can submit. Queue-first admission still accepts valid jobs as queued while queue capacity remains, so a low concurrency number slows throughput without blocking submits.
How should I size a batch with an override?
Use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining, as the budget for new in-flight work. The docs give a worked example: with concurrency_limit: 100, queued_jobs_limit: 500 and nothing running, queue_capacity_remaining is 600 and wave_size_hint is 450. The in-flight budget is still 100, and 450 is only a submission-wave hint that includes queue slots.
Never present wave_size_hint as concurrency, and never substitute plan_concurrency_limit for the effective value. The counts are a snapshot, so they can change right after the response as workers claim jobs or other clients submit.
If you read the limit once and cache it, refresh it from the next submit response, because an override can be added or removed after you cached the number. See pacing bulk submits with generation_limits headroom for the loop.
Sources
Related posts
More in Developers
- curl --retry on a POST: retry a Sume submit with one key
curl --retry also retries a POST, and it resends the same headers each time. Put an Idempotency-Key on a Sume submit first, then pick --retry-max-time.
- Decart lucy-latest vs a pinned Lucy model; Sume catalog ids
Decart's lucy-latest alias can move while legacy Lucy Clip costs $0.15 per second against $0.04 for Lucy 2.5. Why pin a model id, and how to do it on Sume.
- Detect new AI image models: diff Sume GET /v1/images/models
Image models arrive weekly. A short Python diff against GET /v1/images/models tells you when Sume adds or retires an image model id, with no news feed to watch.
- Dub a 20-minute video: audio detach 900 s cap, STT 600 s, TTS 1,200 s
A 20-minute video needs chunking before a dub: audio detach outputs at most 900 s, STT reservation tops out at 600 s, and TTS fails past 1,200 s.
Written by Sume