Sume wave_size_hint and a Worker subrequest limit: submit in waves

A Worker fan-out of Sume jobs hits 50 subrequests on Free. Size each wave from generation_limits, not from the hint alone, and stop at queue_capacity_remaining.

5 min readSume
All posts

Submit Sume jobs from a Worker in waves of at most wave_size_hint, and keep each wave under the Worker subrequest limit, which is 50 on Free and 10,000 on Paid per invocation. The hint is max(1, floor(queue_capacity_remaining * 0.75)). It tells you how many jobs to send now; it is not a concurrency limit and you must not use it to size in-flight work.

Two limits at once

The Cloudflare limits page (updated 2026-09-05) caps subrequests per invocation. Sume caps accepted jobs per workspace. Whichever is smaller sets the wave. A Free-plan Sume workspace accepts 6 jobs in total, so the Sume limit bites first there.

Sume plan capacity and Worker subrequests, read 2026-10-08
PlanProcessingQueueAcceptedWorker subrequests (Free / Paid)
Free15650 / 10,000
Pro4202450 / 10,000
Startup8404850 / 10,000
Scale2010012050 / 10,000

Read the snapshot, not the table

A submit response carries generation_limits with concurrency_limit, active_generation_jobs, queued_generation_jobs, queue_capacity_remaining and wave_size_hint. The counts are a snapshot, and other clients in the same workspace change them. Admin overrides can raise concurrency_limit, so prefer the field over the static plan table.

The in-flight budget is max(0, concurrency_limit - active - queued), capped at queue_capacity_remaining.

def next_wave(limits: dict, already_sent: int = 0) -> int:
    """Jobs to submit now. already_sent counts jobs sent since the snapshot."""
    budget = max(0, limits["concurrency_limit"]
                 - limits["active_generation_jobs"]
                 - limits["queued_generation_jobs"])
    budget = min(budget, limits["queue_capacity_remaining"])
    return max(0, budget - already_sent)

if __name__ == "__main__":
    snap = {"concurrency_limit": 4, "active_generation_jobs": 1,
            "queued_generation_jobs": 0, "queue_capacity_remaining": 23}
    print(next_wave(snap))  # 3

What full looks like

Full concurrency alone is not an error: Sume accepts work as queued while queue capacity remains. When capacity is gone, a paid submit fails with 429 queue_full; request volume over the abuse limit is 429 rate_limited. Retry either with the same Idempotency-Key.

  • Count each submitted job against the budget until the next snapshot.
  • At zero headroom, wait and refresh before you submit.
  • Use a Queue consumer for large batches so each invocation stays inside its subrequest budget.

Worked example for a Pro workspace

On Pro, concurrency is 4, the queue is 20 and the accepted capacity is 24. An empty workspace reports queue_capacity_remaining of 24, so wave_size_hint is floor(24 * 0.75), which is 18. Send 18 jobs, and the snapshot after them shows 4 processing and 14 queued, and remaining capacity of 6. The next hint would be floor(6 * 0.75), which is 4.

The hint leaves a quarter of the room free so that other clients in the same workspace do not hit queue_full. Treat it as a polite submission size. For the pace of work that is actually running, use the budget formula and concurrency_limit.

Staying inside the Worker budget

A Free Worker may make 50 subrequests, so a wave of 18 submits plus polling for them fits only once. Use a Queue and have each consumer invocation send one wave, then stop. A Paid Worker allows 10,000, so the Sume capacity of 120 on Scale is the limit there, not the Worker.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume