asyncio Semaphore size for Sume image batches: accepted capacity

Size the semaphore to what Sume accepts, concurrency plus queue: Free 6, Pro 24, Startup 48, Scale 120. A fake-submit test proves the peak never exceeds it.

4 min readSume
All posts

Sume runs a limited number of generations at once and queues more behind them. The generation admission docs give concurrency of 1 on Free, 4 on Pro, 8 on Startup and 20 on Scale or Enterprise, with queue capacity max(3, 5 x concurrency). Beyond that, a submit is refused with 429 queue_full.

A semaphore sized to concurrency would waste the queue and serialize your batch. A semaphore sized above accepted capacity invites queue_full. The right number is concurrency plus queue.

The numbers

Limits from the Sume generation admission docs, read 2026-10-06
PlanConcurrencyQueue capacityAccepted (semaphore size)
Free156
Pro42024
Startup84048
Scale / Enterprise20100120

Sample with a fake submit

Each task holds the semaphore from submit until the job is terminal. The test submits 40 prompts on the Free size and prints the peak in flight.

import asyncio

ACCEPTED = {"free": 6, "pro": 24, "startup": 48, "scale": 120}  # concurrency + queue

async def run_all(prompts, plan, submit_and_wait):
    gate = asyncio.Semaphore(ACCEPTED[plan])  # accepted capacity, not concurrency

    async def one(prompt):
        async with gate:
            return await submit_and_wait(prompt)

    return await asyncio.gather(*(one(p) for p in prompts))

async def main():
    live = peak = 0
    async def fake(prompt):
        nonlocal live, peak
        live += 1; peak = max(peak, live)
        await asyncio.sleep(0.01)
        live -= 1
        return prompt
    await run_all([f"p{i}" for i in range(40)], "free", fake)
    print("peak in flight:", peak)  # 6, never more than Sume accepts

asyncio.run(main())

Caveats

  • Other pipelines in the same workspace use the same capacity. Read generation_limits and wave_size_hint live when you share a plan.
  • A queue_full refusal releases the idempotency key, so replaying the same key later is safe.
  • Raising your request rate does not raise concurrency; the plan sets it.
  • Wait on jobs with a webhook or waitForJob, not a tight poll.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume