8 active and 15 queued on Startup: the new in-flight budget is zero
Sume's headroom is min(max(0, concurrency - active - queued), queue_capacity_remaining). A Startup example that returns 0, the docs example that returns 60.

For a Startup workspace with 8 jobs processing and 15 queued, the Sume in-flight budget is zero: max(0, 8 - 8 - 15) is 0. The server would still accept more queued work, because queue_capacity_remaining is 25. The budget is a client-side pace control that keeps open work inside the processing cap.
The documented formula
The Generation admission page says to use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs) as the budget for new in-flight work, and to limit that budget to queue_capacity_remaining. Count each newly submitted job against the budget until the next live snapshot. At zero, wait and refresh before you submit more.
| Case | concurrency_limit | active | queued | queue_capacity_remaining | Budget |
|---|---|---|---|---|---|
| Startup, busy | 8 | 8 | 15 | 25 (40 - 15 queued, 0 idle seats) | min(max(0, 8 - 8 - 15), 25) = 0 |
| Docs example | 100 | 30 | 10 | 560 (490 queue + 70 idle seats) | min(max(0, 100 - 30 - 10), 560) = 60 |
The same rule in Python
This function takes a generation_limits object and returns the budget. Run it as is: both cases print without a network call.
def headroom(limits: dict) -> int:
c = limits["concurrency_limit"]
active = limits["active_generation_jobs"]
queued = limits["queued_generation_jobs"]
room = limits["queue_capacity_remaining"]
return min(max(0, c - active - queued), room)
startup = {
"concurrency_limit": 8,
"active_generation_jobs": 8,
"queued_generation_jobs": 15,
"queue_capacity_remaining": 25,
}
docs_example = {
"concurrency_limit": 100,
"active_generation_jobs": 30,
"queued_generation_jobs": 10,
"queue_capacity_remaining": 560,
}
print(headroom(startup)) # 0: wait, refresh, then submit
print(headroom(docs_example)) # 60Why a budget of zero is not an error
Concurrency is a dispatch limit, not a submit limit. A workspace at its processing cap can still receive valid jobs as queued while queue capacity remains. The budget is deliberately stricter than admission, because a long queue delays every job behind it, and cancellation only works before generation starts.
When the budget is zero, poll the jobs you already hold with backoff, and submit again after at least one reaches a terminal state. Do not treat queued as a failure.
Counts are a snapshot
The counts can change immediately after the response, when workers claim jobs or other clients submit. Treat the budget as a conservative guess, not a reservation. If a submit still returns 429 queue_full, wait, then retry with the same idempotency key, as in the queue-full handling section.
Using the budget
The budget is a ceiling on what one wave can add without hitting queue_full, and it already accounts for work that is active or queued. Recompute it from a fresh generation_limits snapshot before each wave rather than carrying the old number forward.
If the budget reads zero, do not sleep for a fixed time. Poll your own jobs, wait for one to reach a terminal state, read the status once more, and refill. A submit that returns 402 insufficient_credits is a different problem: it is a wallet issue, and no amount of waiting changes it.
Sources
Related posts
More in Developers
- Hosted Sume MCP token hygiene: 5 rules for agent credentials
Five Sume credential rules: an OAuth token is not an API key, stays out of CLI config and prompts, never goes to other providers; rotate exposed keys.
- HunyuanVideo 1.5 LoRA training: train.py and the Muon optimizer
HunyuanVideo 1.5 released training code on Dec 5 2025 and says to use the Muon optimizer for LoRA. The flags, the torchrun and FSDP setup, and the hosted gap.
- HunyuanVideo 1.5 cache inference: DeepCache, TeaCache, TaylorCache
HunyuanVideo 1.5 added DeepCache on Nov 24 2025 and TeaCache plus TaylorCache on Nov 27, switched with --enable_cache and --cache_type. What the README claims.
- Hy Image 3.5 multi-turn editing with assembled_history vs Sume edits
How Tencent's Hy Image 3.5 Preview chains edits with assembled_history, and what the same loop looks like on Sume's stateless images route.
Written by Sume