Sume queue capacity: max(3, concurrency x 5), and 7 jobs on Free
Sume's default queue capacity is max(3, concurrency_limit x 5). On Free that is 5 queued plus 1 processing, so a seventh live generation job gets queue_full.

Sume sets default queue capacity to max(3, concurrency_limit × 5). On the Free plan, concurrency is 1 and the queue holds 5, so the accepted job capacity is 6: one job processing and five waiting. A seventh paid generation submit while those six are still live fails with 429 queue_full.
The shape comes from the Generation admission page, and it is worth working through because it explains why a script that submits quickly sees queued first and an error later.
How do concurrency, queue and accepted capacity relate?
Concurrency is a dispatch limit, not a submit limit. A workspace at its concurrency limit can still accept valid jobs as queued while queue capacity remains. The accepted job capacity is concurrency_limit + queued_jobs_limit, the most paid jobs that can be processing or queued for the workspace at once.
The queue only becomes an error when it is full too. Concurrency being full by itself is not an error.
| Plan | Processing | Queue default | Accepted capacity | Check: max(3, c x 5) |
|---|---|---|---|---|
| Free | 1 | 5 | 6 | max(3, 5) = 5 |
| Pro | 4 | 20 | 24 | max(3, 20) = 20 |
| Startup | 8 | 40 | 48 | max(3, 40) = 40 |
| Scale | 20 | 100 | 120 | max(3, 100) = 100 |
What does a seven-job burst on Free look like?
Say your workspace is on Free and idle. You submit seven valid jobs back to back. The first is accepted and starts processing, or sits in queued for a moment before a worker claims it. Jobs two through six are accepted as queued. The seventh finds all six slots taken and the submit fails with 429 queue_full.
The error body can include a generation_limits snapshot for the failed attempt, and Sume releases or refunds the reservation for the failed admission where applicable.
{
"error": {
"code": "queue_full",
"message": "Workspace generation queue is full. Wait for running jobs to complete before submitting more generation work.",
"request_id": "req_..."
}
}Why is the 3 in max(3, ...) there?
The floor means a workspace with a very small concurrency number still gets a minimum queue of 3. With the plan defaults in the table above it never applies, because even Free yields 5. It matters only if an effective concurrency_limit were 0 or 1 under some other configuration, and the docs do not describe such a case, so do not build on it. The table's right column simply shows the formula agrees with the published defaults.
The effective values can differ from the table. Org workspaces have a floor of 10, and admin overrides can raise concurrency. Read generation_limits.concurrency_limit and queued_jobs_limit from a submit response, or the dashboard Concurrency tab, and treat the static table as a default.
What does this mean for a nightly batch?
Plan the batch around accepted capacity rather than around your total job count. On Free, six live jobs is the ceiling, so a 100-image batch has to be released in small groups as earlier jobs finish. On a bigger plan the same logic applies with larger numbers, and the wave_size_hint in the response (max(1, floor(queue_capacity_remaining * 0.75))) is a submission-wave hint only, not a concurrency limit.
Use an idempotency key per item so a retried submit never produces a second paid job. A key such as one built from the batch name and item number is a reasonable convention, and the docs give avatar-batch-001-item-001 as an example of that style.
How do I avoid hitting queue_full?
Do not treat queued as failure, and stop adding work when queue_capacity_remaining is low. Poll existing jobs with backoff until at least one is terminal, cancel queued jobs you no longer need (cancel works only before generation starts), and retry the rejected submit with the same idempotency key. Sume does not expose a precise per-job queue position or ETA today.
For pacing a large wave from the response fields, see pacing bulk submits with generation_limits headroom.
Sources
Related posts
More in Developers
- verifyWebhook returns false during a Sume secret rotation
The npm build of @sume-com/sdk 0.2.0 compares the signature header for equality, so rotation deliveries with two signatures fail. A 21-line fix.
- Does the Sume SDK retry POSTs? Only with an Idempotency-Key
createSumeClient retries 408, 429 and 5xx twice with backoff, but replays a POST only when it carries an Idempotency-Key. How to set it and tune maxRetries.
- @sume-com/sdk waitForJob is not exported: a 26-line replacement
The docs show waitForJob, but the npm build of @sume-com/sdk 0.2.0 does not export it. Here is a fetch version with 429 tolerance and a deadline.
- Check an Avatar payload with sume tools schema before --confirm-paid
sume tools schema avatar-videos.create --json prints the exact fields before you pass --confirm-paid. A read-only pre-flight loop for a spending CLI command.
Written by Sume