Compute Sume job headroom from generation_limits, not a hardcoded 4
Read generation_limits from each Sume submit response and size new work with max(0, concurrency_limit - active - queued). Python example, with the docs numbers.

Compute headroom as max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs), cap it at queue_capacity_remaining, and read all three numbers from the generation_limits object in the latest submit response instead of writing a constant such as 4 into your code. Sume's generation admission docs use this formula: with a limit of 100, 30 processing and 10 queued jobs, the new in-flight budget is 60.
A constant is wrong the day the plan changes, and it is wrong when an admin override applies. limit_source tells you which one you have, and concurrency_limit is the effective value in both cases. Do not use plan_concurrency_limit or wave_size_hint for this sum: the docs say the hint is a submission-wave hint and never a width for in-flight work.
A headroom function
The function below takes the snapshot dict as the API returns it, so you can call it with the generation_limits of your last submit. It runs offline with the numbers from the docs.
import asyncio
def headroom(gl: dict) -> int:
"""New in-flight jobs to submit now, from a generation_limits snapshot."""
limit = int(gl["concurrency_limit"])
active = int(gl.get("active_generation_jobs", 0))
queued = int(gl.get("queued_generation_jobs", 0))
room = max(0, limit - active - queued)
return min(room, int(gl.get("queue_capacity_remaining", room)))
async def main() -> None:
docs_example = {
"limit_source": "plan",
"concurrency_limit": 100,
"active_generation_jobs": 30,
"queued_generation_jobs": 10,
"queue_capacity_remaining": 560,
"wave_size_hint": 420,
}
print("headroom:", headroom(docs_example))
print("full:", headroom({**docs_example, "active_generation_jobs": 100}))
asyncio.run(main())How to read each field
The counts are a snapshot. Workers claim jobs and other clients submit between your response and your next call, so treat zero as "wait" and any positive number as an upper bound.
| Field | Use it for | Do not use it for |
|---|---|---|
| concurrency_limit | Effective processing seats; the sum above | Hardcoding a plan default |
| active_generation_jobs | Jobs with status processing | Counting only your own jobs |
| queued_generation_jobs | Jobs with status queued | Ignoring other clients in the workspace |
| queue_capacity_remaining | Hard cap before queue_full | Sizing processing width |
| wave_size_hint | Size of a submission wave | Sizing in-flight work |
| limit_source | Telling plan from admin_override | Branching on a plan name |
What happens at zero
At zero headroom, refresh the preview and wait before you submit more. If you do submit past the queue, the API answers 429 queue_full, a different code from rate_limited: queue_full means the workspace has no queue capacity, so wait for jobs to finish, not for a timer. The two codes need different handling, and a backoff loop that treats them alike either hammers a full queue or sleeps while capacity is free.
Sources
Related posts
More in Developers
- Recraft V4.1 Flash: median 1.3 s, p95 1.8 s. Set timeouts from p95
Recraft quotes a median of about 1.3 seconds and a p95 of 1.8 seconds for V4.1 Flash. How to turn latency claims into timeouts and polling for image APIs.
- Recraft V4 on Sume returns WebP only: convert to PNG with Pillow
Recraft V4 on Sume lists webp as its only output format at $0.05 per image. Download the file and convert to PNG or JPEG with Pillow in six lines.
- Redeliver a missed speech-to-text webhook without rerunning the job
Your receiver was down when a Sume STT job finished. Redeliver the terminal webhook with one call instead of paying to transcribe again.
- Reels safe-zone boxes in Python: a pure function for any frame size
Meta's Reels percentages are relative, so one pure Python function gives the safe box for any frame size. It runs offline and feeds a caption anchor.
Written by Sume