generation_limits missing from a Sume submit: send one, then re-read
The docs say the snapshot is included when Sume can compute it. A Python rule for waves when it is absent, full, or open, tested on three inputs.

Do not assume a Sume submit response always carries generation_limits. The docs say it is included "when Sume can compute the workspace admission snapshot", and they tell you to refresh before you choose a width when the counts are not available. A safe rule: if the object is missing, send one job, read again, and only then widen the wave.
What the docs promise, and what they do not
The Generation admission page describes the object and its fields, then says the counts are a snapshot that can change right after the response. It does not promise that every response has one. A client that indexes response['generation_limits']['concurrency_limit'] without a check will raise a KeyError on the day it is absent.
| Snapshot | What to do | Why |
|---|---|---|
| Missing | Submit one job, then re-read | The docs say to refresh before you select a width |
| Present, headroom 0 | Wait, poll your own jobs, refresh | At zero headroom, wait and refresh the preview |
| Present, headroom above 0 | Submit up to the headroom, capped by queue_capacity_remaining | Count each new job against the budget until the next snapshot |
The function
The function returns how many of your planned jobs to send now. It reuses the headroom formula from the docs. I ran the three cases offline.
def room_for_new_jobs(submit_response: dict, planned: int) -> int:
"""How many of `planned` jobs to send now. 0 means: refresh first."""
limits = submit_response.get("generation_limits")
if not limits:
return 1 if planned else 0 # snapshot missing: one probe job, then re-read
c = limits["concurrency_limit"]
room = min(
max(0, c - limits["active_generation_jobs"] - limits["queued_generation_jobs"]),
limits["queue_capacity_remaining"],
)
return min(planned, room)
full = {"generation_limits": {"concurrency_limit": 4, "active_generation_jobs": 4,
"queued_generation_jobs": 20, "queue_capacity_remaining": 0}}
open_ = {"generation_limits": {"concurrency_limit": 4, "active_generation_jobs": 1,
"queued_generation_jobs": 0, "queue_capacity_remaining": 23}}
print(room_for_new_jobs({}, 10)) # 1
print(room_for_new_jobs(full, 10)) # 0
print(room_for_new_jobs(open_, 10)) # 3Keep the idempotency key stable
When you send the single probe job and then wait, give each job its own Idempotency-Key and reuse that key if you have to resend. A retry with the same key and the same payload returns the original job and does not bill a second one. A different payload under the same key returns 409 idempotency_conflict.
If the probe returns 429 queue_full, treat that as the answer: wait for jobs to finish or cancel queued ones, then retry with the same key.
Making the probe cheap
The probe job should be a real piece of work you need anyway, such as the first item of the batch, not a throwaway. After it returns, read the snapshot, compute the headroom and send the rest in waves.
If the snapshot never appears on your workspace, fall back to a fixed small wave and rely on 429 queue_full with a backoff. That costs a few rejected writes but keeps the pipeline moving.
Sources
Related posts
More in Developers
- HappyHorse 1.1 [Image 1] tags vs Sume's <IMAGE_REF_0> in prompts
HappyHorse reference-to-video counts images from [Image 1]; Sume's Omni counts from <IMAGE_REF_0>. How to map a prompt without an off-by-one.
- 8 active and 15 queued on Startup: the new in-flight budget is zero
Sume's headroom is min(max(0, concurrency - active - queued), queue_capacity_remaining). A Startup example that returns 0, the docs example that returns 60.
- Hosted Sume MCP token hygiene: 5 rules for agent credentials
Five Sume credential rules: an OAuth token is not an API key, stays out of CLI config and prompts, never goes to other providers; rotate exposed keys.
- HunyuanVideo 1.5 LoRA training: train.py and the Muon optimizer
HunyuanVideo 1.5 released training code on Dec 5 2025 and says to use the Muon optimizer for LoRA. The flags, the torchrun and FSDP setup, and the hosted gap.
Written by Sume