wave_size_hint 450 is not concurrency: sizing Sume submit waves
With concurrency 100 and a 500-job queue, Sume shows wave_size_hint 450 but your in-flight budget is 100. How to compute both from generation_limits.

wave_size_hint is a hint for how many jobs to submit in one wave, and it counts queue slots. It is not your concurrency. With a workspace set to 100 concurrent jobs and 500 queue slots and nothing running, Sume reports queue_capacity_remaining 600 and wave_size_hint 450, but the room for new in-flight work is 100. Compute the in-flight budget as max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs) and cap it at queue_capacity_remaining.
The fields, with the worked example
Submit responses include a generation_limits snapshot. The counts can change right after the response, because workers claim jobs and other clients submit work.
| Field or result | Idle workspace | 30 processing, 10 queued |
|---|---|---|
| concurrency_limit | 100 | 100 |
| queued_jobs_limit | 500 | 500 |
| queue_capacity_remaining | 600 | 490 queue + 70 idle seats = 560 |
| wave_size_hint = max(1, floor(remaining x 0.75)) | 450 | floor(560 x 0.75) = 420 |
| New in-flight budget | 100 | 100 - 30 - 10 = 60 |
Plan defaults for comparison
On the plan defaults, the default queue capacity is max(3, concurrency_limit x 5). Always prefer the effective generation_limits fields over a static table, because admin overrides can raise the cap.
| Plan | Processing | Queue | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
A sizing function
The function returns how many jobs to submit next. At zero headroom it returns 0, and you should wait and refresh the snapshot. Do not use plan_concurrency_limit, which ignores overrides.
def next_wave(limits: dict) -> int:
conc = limits["concurrency_limit"]
active = limits.get("active_generation_jobs", 0)
queued = limits.get("queued_generation_jobs", 0)
remaining = limits.get("queue_capacity_remaining", 0)
headroom = max(0, conc - active - queued)
return min(headroom, remaining)
idle = {"concurrency_limit": 100, "queued_jobs_limit": 500,
"active_generation_jobs": 0, "queued_generation_jobs": 0,
"queue_capacity_remaining": 600}
busy = dict(idle, active_generation_jobs=30, queued_generation_jobs=10,
queue_capacity_remaining=560)
print(next_wave(idle)) # 100
print(next_wave(busy)) # 60What queue_full means
If both processing and queue are full, a submit gets 429 queue_full. Wait for a job to finish or cancel queued ones, then retry with the same idempotency key. A full processing cap alone is not an error, because accepted jobs wait as queued.
Why the hint exists at all
The docs define the hint as max(1, floor(queue_capacity_remaining x 0.75)) and call it a submission-wave hint only, not a concurrency limit, an override or a processing width. Jobs in the queue still wait for a processing seat, and only concurrency_limit of them run at a time.
So two numbers answer two different questions. The in-flight budget answers how many jobs will run soon. The wave hint answers how many you can hand over in one burst without risking queue_full for yourself or a teammate.
The counts are a snapshot. After each wave, read a fresh generation_limits from the next submit response or from your own accounting, and subtract what you have just sent. At zero headroom, wait. Never present the hint as your concurrency, and never compute from plan_concurrency_limit when an admin override applies. The effective concurrency_limit field already includes it.
Sources
Related posts
More in Developers
- Webhook signature mismatch: Sume fingerprint vs OpenRouter t=,v1=
A video webhook that fails verification has three usual causes: parsed body, wrong secret, wrong header format. Sume adds a secret fingerprint header.
- Check your Sume webhook secret at boot: 64 hex and a match
A Node startup check for a Sume webhook receiver: the secret must be 64 hex characters, non-empty, and equal to GET /v1/webhooks/signing-secret.
- Test a video webhook receiver before the first job: Sume vs fal
Sume sends a signed dummy webhook.test with one API call. fal retries real results up to 31 times and treats 3xx as failure, so test your URL first.
- Which key scope each Sume webhook endpoint needs
Reading the signing secret needs account:read, rotating and Send test need account:write, and redelivering a job needs jobs:write. A least-privilege map.
Written by Sume