wave_size_hint 450 is not concurrency: sizing Sume submit waves

With concurrency 100 and a 500-job queue, Sume shows wave_size_hint 450 but your in-flight budget is 100. How to compute both from generation_limits.

5 min readSume
All posts

wave_size_hint is a hint for how many jobs to submit in one wave, and it counts queue slots. It is not your concurrency. With a workspace set to 100 concurrent jobs and 500 queue slots and nothing running, Sume reports queue_capacity_remaining 600 and wave_size_hint 450, but the room for new in-flight work is 100. Compute the in-flight budget as max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs) and cap it at queue_capacity_remaining.

The fields, with the worked example

Submit responses include a generation_limits snapshot. The counts can change right after the response, because workers claim jobs and other clients submit work.

generation_limits worked example from the Sume docs, as of 2026-10-08
Field or resultIdle workspace30 processing, 10 queued
concurrency_limit100100
queued_jobs_limit500500
queue_capacity_remaining600490 queue + 70 idle seats = 560
wave_size_hint = max(1, floor(remaining x 0.75))450floor(560 x 0.75) = 420
New in-flight budget100100 - 30 - 10 = 60

Plan defaults for comparison

On the plan defaults, the default queue capacity is max(3, concurrency_limit x 5). Always prefer the effective generation_limits fields over a static table, because admin overrides can raise the cap.

Plan concurrency and queue defaults, as of 2026-10-08
PlanProcessingQueueAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120

A sizing function

The function returns how many jobs to submit next. At zero headroom it returns 0, and you should wait and refresh the snapshot. Do not use plan_concurrency_limit, which ignores overrides.

def next_wave(limits: dict) -> int:
    conc = limits["concurrency_limit"]
    active = limits.get("active_generation_jobs", 0)
    queued = limits.get("queued_generation_jobs", 0)
    remaining = limits.get("queue_capacity_remaining", 0)
    headroom = max(0, conc - active - queued)
    return min(headroom, remaining)

idle = {"concurrency_limit": 100, "queued_jobs_limit": 500,
        "active_generation_jobs": 0, "queued_generation_jobs": 0,
        "queue_capacity_remaining": 600}
busy = dict(idle, active_generation_jobs=30, queued_generation_jobs=10,
            queue_capacity_remaining=560)
print(next_wave(idle))  # 100
print(next_wave(busy))  # 60

What queue_full means

If both processing and queue are full, a submit gets 429 queue_full. Wait for a job to finish or cancel queued ones, then retry with the same idempotency key. A full processing cap alone is not an error, because accepted jobs wait as queued.

Why the hint exists at all

The docs define the hint as max(1, floor(queue_capacity_remaining x 0.75)) and call it a submission-wave hint only, not a concurrency limit, an override or a processing width. Jobs in the queue still wait for a processing seat, and only concurrency_limit of them run at a time.

So two numbers answer two different questions. The in-flight budget answers how many jobs will run soon. The wave hint answers how many you can hand over in one burst without risking queue_full for yourself or a teammate.

The counts are a snapshot. After each wave, read a fresh generation_limits from the next submit response or from your own accounting, and subtract what you have just sent. At zero headroom, wait. Never present the hint as your concurrency, and never compute from plan_concurrency_limit when an admin override applies. The effective concurrency_limit field already includes it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume