asyncio Semaphore size for Sume image batches: accepted capacity
Size the semaphore to what Sume accepts, concurrency plus queue: Free 6, Pro 24, Startup 48, Scale 120. A fake-submit test proves the peak never exceeds it.

Sume runs a limited number of generations at once and queues more behind them. The generation admission docs give concurrency of 1 on Free, 4 on Pro, 8 on Startup and 20 on Scale or Enterprise, with queue capacity max(3, 5 x concurrency). Beyond that, a submit is refused with 429 queue_full.
A semaphore sized to concurrency would waste the queue and serialize your batch. A semaphore sized above accepted capacity invites queue_full. The right number is concurrency plus queue.
The numbers
| Plan | Concurrency | Queue capacity | Accepted (semaphore size) |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale / Enterprise | 20 | 100 | 120 |
Sample with a fake submit
Each task holds the semaphore from submit until the job is terminal. The test submits 40 prompts on the Free size and prints the peak in flight.
import asyncio
ACCEPTED = {"free": 6, "pro": 24, "startup": 48, "scale": 120} # concurrency + queue
async def run_all(prompts, plan, submit_and_wait):
gate = asyncio.Semaphore(ACCEPTED[plan]) # accepted capacity, not concurrency
async def one(prompt):
async with gate:
return await submit_and_wait(prompt)
return await asyncio.gather(*(one(p) for p in prompts))
async def main():
live = peak = 0
async def fake(prompt):
nonlocal live, peak
live += 1; peak = max(peak, live)
await asyncio.sleep(0.01)
live -= 1
return prompt
await run_all([f"p{i}" for i in range(40)], "free", fake)
print("peak in flight:", peak) # 6, never more than Sume accepts
asyncio.run(main())Caveats
- Other pipelines in the same workspace use the same capacity. Read
generation_limitsandwave_size_hintlive when you share a plan. - A
queue_fullrefusal releases the idempotency key, so replaying the same key later is safe. - Raising your request rate does not raise concurrency; the plan sets it.
- Wait on jobs with a webhook or
waitForJob, not a tight poll.
Sources
Related posts
More in Developers
- asyncio TaskGroup cancels siblings: poll many Sume jobs safely
A TaskGroup cancels every other poller when one raises. For a batch of Sume video jobs that abandons waits, not jobs. Catch inside the task and return results.
- Check balance before bulk transcription: 402 and admission preview
A 402 insufficient_credits arrives before any provider work. Read GET /v1/balance and POST /v1/generation/admission-preview first and size the run.
- Port a Bedrock image call to Sume /v1/images in Python
Moving from boto3 invoke_model for Nova Canvas or Titan to Sume's REST call: the request mapping, the response shape, the status codes, and the swap code.
- BullMQ delayed job that polls an AI video job and reschedules itself
A BullMQ worker reads Sume's job status once, then adds the next poll with a delay from next_poll_after_seconds, so no worker slot is held while a clip renders.
Written by Sume