Why concurrency_limit differs from the Sume plan table

In Sume generation_limits, concurrency_limit is the effective cap and limit_source says plan or admin_override. Size work from it, not plan_concurrency_limit.

5 min readSume
All posts

If the generation_limits block on a Sume submit response shows a concurrency_limit that differs from the plan table in the docs, trust concurrency_limit. It is the effective cap, and limit_source tells you where it came from: plan for the plan default, admin_override when Sume raised it for the workspace.

The field next to it, plan_concurrency_limit, is only the plan default. The Generation admission page says to ignore it for wave sizing when an override applies.

Which fields tell me the real limit?

A submit response includes a snapshot like this when Sume can compute it:

  • concurrency_limit is the effective maximum number of same-workspace paid generation jobs that can be processing at once.
  • limit_source is plan or admin_override.
  • plan_concurrency_limit is the plan default and can differ from the effective value.
  • accepted_generation_jobs_limit is concurrency_limit + queued_jobs_limit.
{
  "generation_limits": {
    "plan_id": "pro",
    "limit_source": "plan",
    "plan_concurrency_limit": 4,
    "concurrency_limit": 4,
    "queued_jobs_limit": 20,
    "accepted_generation_jobs_limit": 24
  }
}

What are the plan defaults, and why do they not always match?

The docs list static defaults per plan. They are a starting point, and the page itself says to prefer the effective field over the table.

Three things can move your effective number. An admin override raises it (limit_source: admin_override). Org workspaces have a floor of 10. And Enterprise defaults to 20 and uses admin overrides for higher contract limits.

Default generation limits by plan, from the Sume docs (read 2026-10-02)
PlanProcessing concurrencyQueue capacity (default)Accepted job capacity
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

Do top-ups raise my concurrency?

No. Generation concurrency is plan-only. Prepaid top-ups do not raise the processing concurrency limit, so adding funds will not change concurrency_limit. The dashboard Concurrency tab is the source of truth for the workspace's configured cap, exposed as generation_limits.concurrency_limit.

Queue capacity defaults to max(3, concurrency_limit × 5). With an override that raises concurrency, the default queue grows with it, and queued_jobs_limit in the response shows the figure that applies to you.

Why does my dashboard number not match this post?

Because the post quotes defaults. The Concurrency tab shows your workspace's configured processing cap, and the API exposes the same figure as generation_limits.concurrency_limit. If the two ever disagree, the docs name the API field as the effective one.

Also keep purchased_concurrency_limit-style fields out of your sizing logic. The docs say never to substitute the pre-override fields for the effective one, and top-ups do not change processing concurrency at all.

Finally, remember what the cap controls. It limits how many generation jobs are processing at once, not how many you can submit. Queue-first admission still accepts valid jobs as queued while queue capacity remains, so a low concurrency number slows throughput without blocking submits.

How should I size a batch with an override?

Use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), capped by queue_capacity_remaining, as the budget for new in-flight work. The docs give a worked example: with concurrency_limit: 100, queued_jobs_limit: 500 and nothing running, queue_capacity_remaining is 600 and wave_size_hint is 450. The in-flight budget is still 100, and 450 is only a submission-wave hint that includes queue slots.

Never present wave_size_hint as concurrency, and never substitute plan_concurrency_limit for the effective value. The counts are a snapshot, so they can change right after the response as workers claim jobs or other clients submit.

If you read the limit once and cache it, refresh it from the next submit response, because an override can be added or removed after you cached the number. See pacing bulk submits with generation_limits headroom for the loop.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume