Compute Sume job headroom from generation_limits, not a hardcoded 4

Read generation_limits from each Sume submit response and size new work with max(0, concurrency_limit - active - queued). Python example, with the docs numbers.

5 min readSume
All posts

Compute headroom as max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), cap it at queue_capacity_remaining, and read all three numbers from the generation_limits object in the latest submit response instead of writing a constant such as 4 into your code. Sume's generation admission docs use this formula: with a limit of 100, 30 processing and 10 queued jobs, the new in-flight budget is 60.

A constant is wrong the day the plan changes, and it is wrong when an admin override applies. limit_source tells you which one you have, and concurrency_limit is the effective value in both cases. Do not use plan_concurrency_limit or wave_size_hint for this sum: the docs say the hint is a submission-wave hint and never a width for in-flight work.

A headroom function

The function below takes the snapshot dict as the API returns it, so you can call it with the generation_limits of your last submit. It runs offline with the numbers from the docs.

import asyncio

def headroom(gl: dict) -> int:
    """New in-flight jobs to submit now, from a generation_limits snapshot."""
    limit = int(gl["concurrency_limit"])
    active = int(gl.get("active_generation_jobs", 0))
    queued = int(gl.get("queued_generation_jobs", 0))
    room = max(0, limit - active - queued)
    return min(room, int(gl.get("queue_capacity_remaining", room)))

async def main() -> None:
    docs_example = {
        "limit_source": "plan",
        "concurrency_limit": 100,
        "active_generation_jobs": 30,
        "queued_generation_jobs": 10,
        "queue_capacity_remaining": 560,
        "wave_size_hint": 420,
    }
    print("headroom:", headroom(docs_example))
    print("full:", headroom({**docs_example, "active_generation_jobs": 100}))

asyncio.run(main())

How to read each field

The counts are a snapshot. Workers claim jobs and other clients submit between your response and your next call, so treat zero as "wait" and any positive number as an upper bound.

generation_limits fields for sizing work (Sume docs, read 2026-10-04)
FieldUse it forDo not use it for
concurrency_limitEffective processing seats; the sum aboveHardcoding a plan default
active_generation_jobsJobs with status processingCounting only your own jobs
queued_generation_jobsJobs with status queuedIgnoring other clients in the workspace
queue_capacity_remainingHard cap before queue_fullSizing processing width
wave_size_hintSize of a submission waveSizing in-flight work
limit_sourceTelling plan from admin_overrideBranching on a plan name

What happens at zero

At zero headroom, refresh the preview and wait before you submit more. If you do submit past the queue, the API answers 429 queue_full, a different code from rate_limited: queue_full means the workspace has no queue capacity, so wait for jobs to finish, not for a timer. The two codes need different handling, and a backoff loop that treats them alike either hammers a full queue or sleeps while capacity is free.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume