asyncio Semaphore: limit concurrent API jobs in Python

An asyncio Semaphore caps how many coroutines run a block at once. For paid API jobs, hold it from submit to the final status and size it to your limit.

5 min readSume
All posts

An asyncio.Semaphore(n) limits how many coroutines can be inside an async with sem: block at the same time: each entry decrements an internal counter, and when the counter is zero the next coroutine waits until another one leaves. To limit concurrent API jobs, hold the semaphore for each job's whole life, from the submit to its final status, not just for the HTTP request.

The asyncio facts below come from the Python documentation, read 2026-09-28. The job API in the example is Sume's; its limits come from the generation admission docs.

How does asyncio.Semaphore work?

Python's docs: "A semaphore manages an internal counter which is decremented by each acquire() call and incremented by each release() call." The counter never goes below zero; when acquire() finds it at zero, it blocks until some task calls release(). The starting value defaults to 1, and "the preferred way to use a Semaphore is an async with statement", which releases even when the block raises (asyncio).

Prefer asyncio.BoundedSemaphore: it raises ValueError if release() would push the counter above its starting value, which turns a double release into an error instead of a silent extra slot. Neither class is thread-safe, so use them inside one event loop.

How do I limit concurrent API jobs with a semaphore?

Put the submit and the polling inside the same async with block. The example uses HTTPX's AsyncClient (HTTPX) against POST /v1/videos, which returns a job with a polling_url at once (Video generation). Each item gets its own stable Idempotency-Key, so rerunning the script with the same prompts replays accepted jobs instead of paying twice.

import asyncio, os, httpx

API = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
DONE = ("completed", "failed", "cancelled")
sem = asyncio.BoundedSemaphore(4)  # your workspace's concurrency_limit

async def make_video(client, item_id, prompt):
    async with sem:  # held from submit to final status
        r = await client.post(API, headers={**AUTH, "Idempotency-Key": f"batch-7-{item_id}"},
                              json={"model": "seedance-2", "prompt": prompt, "resolution": "720p"})
        r.raise_for_status()
        job = r.json()
        while job["status"] not in DONE:
            await asyncio.sleep(30)
            job = (await client.get(job["polling_url"], headers=AUTH)).json()
        return job

async def main(prompts):
    async with httpx.AsyncClient() as client:
        return await asyncio.gather(*(make_video(client, i, p) for i, p in enumerate(prompts)))

Why hold the semaphore until the job finishes?

Because the POST returns in a moment, and the job runs for minutes. If the semaphore wraps only the request, a 500-prompt batch is submitted within seconds, and the semaphore limits nothing that matters.

Sume accepts jobs beyond your processing limit as queued, since "concurrency is a dispatch limit, not a submit limit". Every accepted job reserves its estimated cost at submit, and once the queue is also full, new submits fail with 429 queue_full (Generation admission). Holding the slot until the final status keeps in-flight work, and the balance reserved for it, at your processing limit.

What number should the semaphore be?

Start from your workspace's effective concurrency_limit; the docs call the dashboard Concurrency tab its source of truth (Generation admission). Video job concurrency and queueing shows how to size from a live generation_limits snapshot when other processes share the workspace.

A semaphore only counts the coroutines in its own event loop. Two scripts that each create BoundedSemaphore(4) can have eight jobs in flight against the same workspace, so split the limit between them.

Polling spends the read budget, which is separate from the write budget that submits spend, so a status loop cannot 429 your own submits.

What should each job do when a submit fails?

Decide per error code, inside the semaphore, so a failing job gives its slot back. A client-side timeout does not cancel a job, so never resubmit one just because your coroutine gave up.

From Sume's Generation admission and Jobs and results docs, read 2026-09-28; the queue_full replay is current code.
ResponseMeaningWhat the coroutine should do
429 queue_fullNo accepted-job capacity leftStop adding work, wait for a job to finish, then submit under a new key; in current code a same-key resend returns the same queue_full
429 rate_limitedRequest budget exceededWait for retry-after, then retry with the same key
402 insufficient_creditsBalance can't cover the estimateStop the batch
409 idempotency_conflictKey reused with a different bodyFix the key scheme; reuse keys only for exact retries

Sources

Related posts

More in Developers

All Developers posts

Written by Sume