Python asyncio.Semaphore sized to a Sume plan's job capacity
Free holds 6 paid jobs at once, Pro 24, Startup 48, Scale 120. A Semaphore of that size keeps a 60-job batch from hitting 429 queue_full.

Set the semaphore to concurrency plus queue capacity for your plan, and release a slot only when a job reaches a terminal state. Sume accepts a valid paid job as queued when workers are busy, so concurrency alone is not the submit limit. The submit limit is accepted capacity, and a submit beyond it fails with 429 queue_full.
The numbers
Sume's admission docs give default queue capacity as max(3, concurrency_limit x 5). The docs also say to prefer the effective generation_limits fields in a submit response over this static table, because admin overrides change them.
| Plan | Processing | Queue | Accepted |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
A runnable shape
The sample fakes the submit with a short sleep, so it runs offline. It asserts nothing about Sume. It shows the part you own: sixty tasks start, and at most 24 hold a slot at once on Pro. In a real worker, keep the slot through polling, since a queued job still counts against capacity.
import asyncio
PLANS = {"free": (1, 5), "pro": (4, 20), "startup": (8, 40), "scale": (20, 100)} # (concurrency, queue)
async def submit(i): # stand-in for your POST with an Idempotency-Key
await asyncio.sleep(0.01)
return f"job_{i}"
async def main(plan="pro", total=60):
concurrency, queue = PLANS[plan]
accepted = concurrency + queue # jobs Sume holds at once; the next submit gets 429 queue_full
gate = asyncio.Semaphore(accepted)
inflight = 0
peak = 0
async def one(i):
nonlocal inflight, peak
async with gate: # in a real worker, release when the job turns terminal, not after submit
inflight += 1
peak = max(peak, inflight)
job = await submit(i)
await asyncio.sleep(0.05)
inflight -= 1
return job
jobs = await asyncio.gather(*(one(i) for i in range(total)))
print(len(jobs), "submitted, peak", peak, "of", accepted)
asyncio.run(main())Caveats
Capacity is per workspace, not per process. Two workers each holding 24 slots can still collide, so share the budget or read queue_capacity_remaining from generation_limits before each wave. wave_size_hint in the same object is only a hint, not the concurrency limit.
A 429 can also be rate_limited, a request-volume limit with its own retry-after. Handle the two codes separately: one means wait for jobs to finish, the other means slow down.
Sources
Related posts
More in Developers
- Python: check Omni Flash 1.1 limits against /v1/videos/models first
Google's Omni runs a sync call and extends in 10-second steps up to 40 seconds. Sume's gemini-omni-flash-1.1 takes 3 to 10 seconds per job. A preflight check.
- Python CI guard: fail the build on retired Veo, Imagen, Sora ids
Twenty-three lines of standard-library Python scan a repo for Veo 2/3, Imagen 4 and sora-2 ids with their shutdown dates, and exit 1 when one is found.
- Python guard: refuse a Short over 180 s before the render bills
A 20-line Python check that blocks a Shorts render over 180 seconds and prices the Sume Timeline minutes (10 cents each) before you submit.
- Python: cheapest Sume image model that lists your aspect ratio
A 25-line Python script reads Sume's image catalog, keeps models that list your aspect ratio, prices each from its endpoints record and prints the cheapest.
Written by Sume