generation_limits missing from a Sume submit: send one, then re-read

The docs say the snapshot is included when Sume can compute it. A Python rule for waves when it is absent, full, or open, tested on three inputs.

4 min readSume
All posts

Do not assume a Sume submit response always carries generation_limits. The docs say it is included "when Sume can compute the workspace admission snapshot", and they tell you to refresh before you choose a width when the counts are not available. A safe rule: if the object is missing, send one job, read again, and only then widen the wave.

What the docs promise, and what they do not

The Generation admission page describes the object and its fields, then says the counts are a snapshot that can change right after the response. It does not promise that every response has one. A client that indexes response['generation_limits']['concurrency_limit'] without a check will raise a KeyError on the day it is absent.

Client decision by snapshot state (Sume docs, read 2026-10-09)
SnapshotWhat to doWhy
MissingSubmit one job, then re-readThe docs say to refresh before you select a width
Present, headroom 0Wait, poll your own jobs, refreshAt zero headroom, wait and refresh the preview
Present, headroom above 0Submit up to the headroom, capped by queue_capacity_remainingCount each new job against the budget until the next snapshot

The function

The function returns how many of your planned jobs to send now. It reuses the headroom formula from the docs. I ran the three cases offline.

def room_for_new_jobs(submit_response: dict, planned: int) -> int:
    """How many of `planned` jobs to send now. 0 means: refresh first."""
    limits = submit_response.get("generation_limits")
    if not limits:
        return 1 if planned else 0  # snapshot missing: one probe job, then re-read
    c = limits["concurrency_limit"]
    room = min(
        max(0, c - limits["active_generation_jobs"] - limits["queued_generation_jobs"]),
        limits["queue_capacity_remaining"],
    )
    return min(planned, room)

full = {"generation_limits": {"concurrency_limit": 4, "active_generation_jobs": 4,
        "queued_generation_jobs": 20, "queue_capacity_remaining": 0}}
open_ = {"generation_limits": {"concurrency_limit": 4, "active_generation_jobs": 1,
         "queued_generation_jobs": 0, "queue_capacity_remaining": 23}}
print(room_for_new_jobs({}, 10))     # 1
print(room_for_new_jobs(full, 10))   # 0
print(room_for_new_jobs(open_, 10))  # 3

Keep the idempotency key stable

When you send the single probe job and then wait, give each job its own Idempotency-Key and reuse that key if you have to resend. A retry with the same key and the same payload returns the original job and does not bill a second one. A different payload under the same key returns 409 idempotency_conflict.

If the probe returns 429 queue_full, treat that as the answer: wait for jobs to finish or cancel queued ones, then retry with the same key.

Making the probe cheap

The probe job should be a real piece of work you need anyway, such as the first item of the batch, not a throwaway. After it returns, read the snapshot, compute the headroom and send the rest in waves.

If the snapshot never appears on your workspace, fall back to a fixed small wave and rely on 429 queue_full with a backoff. That costs a few rejected writes but keeps the pipeline moving.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume