Org workspace concurrency floor of 10: the queue and wave that follow
Sume gives org workspaces a processing floor of 10. With the default queue formula that is 50 queued, 60 accepted, and a wave hint of 45. Read your own fields.

Sume docs say that org workspaces have a processing concurrency floor of 10. If the default queue formula applies to that effective limit, the queue holds 50 jobs, the accepted capacity is 60, and an empty workspace shows a wave_size_hint of 45. Treat that as arithmetic from the documented formulas, and confirm it against generation_limits in a real submit response.
Where the floor comes from
The Generation admission page lists the plan defaults: Free 1, Pro 4, Startup 8, Scale 20, Enterprise 20. It then adds that org workspaces have a floor of 10. A Startup org workspace, which would default to 8, is therefore described as running with at least 10. Enterprise has a default of 20, and higher contract limits use admin overrides.
Concurrency is plan-only. Prepaid top-ups do not raise it. Admin overrides can raise the effective concurrency_limit, and the response then reports limit_source: admin_override.
The arithmetic
The docs define the default queue capacity as max(3, concurrency_limit x 5). They define accepted capacity as concurrency plus queue, and the hint as max(1, floor(queue_capacity_remaining * 0.75)).
| Field | Formula | Value |
|---|---|---|
| concurrency_limit | org floor | 10 |
| queued_jobs_limit | max(3, 10 x 5) | 50 |
| accepted_generation_jobs_limit | 10 + 50 | 60 |
| queue_capacity_remaining | 60 - 0 queued - 0 active | 60 |
| wave_size_hint | floor(60 x 0.75) | 45 |
What to hard-code, and what not to
Do not hard-code 10, 50, or 60. The docs tell you to prefer the effective field over their static table, and an override changes every derived number. Read concurrency_limit, queued_jobs_limit, and queue_capacity_remaining from generation_limits on each submit response.
Do not use plan_concurrency_limit to size work when limit_source is admin_override. It is the plan default, not the effective cap. The override post shows the two fields side by side.
A quick way to see your own numbers
Submit one cheap, valid job and read the generation_limits object from the response. The counts are a snapshot and change as workers claim jobs and as other clients submit. If the object is missing, Sume could not compute the snapshot for that response, and the docs say to refresh before you pick a width.
Where the floor shows up
The floor matters when you size a worker pool. A pool sized to the Pro value of 4 would leave capacity unused on an org workspace, and a pool sized to 60 accepted jobs would be right. The figures above are derived from the documented formulas, not read from a live workspace, so confirm them with the generation_limits object on a real submit response.
Concurrency is a dispatch limit. A submit beyond it is not rejected; it waits in the queue. The rejection is 429 queue_full, and it happens only when concurrency plus queue is exhausted. That is why pacing by accepted capacity, not by concurrency, is the pattern that avoids both idle workers and 429s.
Sources
Related posts
More in Developers
- Pick a Sume video model in code: audio refs, 1080p, 20 seconds
Filter GET /v1/video-router/models on reference_audios, resolutions and duration_seconds, then submit the survivor with reference_audio_urls (up to five).
- Pick the Sume video model in code: seconds, ratio, edit or swap
A 12-line Python function turns duration, aspect ratio, edit and swap needs into the models that can run the job. A 4 s 1:1 clip has none; a 12 s one has two.
- Poll every 2 s or 30 s? Reads per 20-minute video job on Free and Pro
Polling a Sume video job every 2 s is 600 reads in 20 minutes; every 30 s is 40. Share of the read budget on Free and Pro, with a Python check.
- Status reads for a full Sume queue: 3,600 to 72,000
Poll every accepted job every 2 s for 20 minutes: Free makes 3,600 reads, Scale 72,000. The math, and how one field cuts it.
Written by Sume