asyncio Semaphore: limit concurrent API jobs in Python
An asyncio Semaphore caps how many coroutines run a block at once. For paid API jobs, hold it from submit to the final status and size it to your limit.

An asyncio.Semaphore(n) limits how many coroutines can be inside an async with sem: block at the same time: each entry decrements an internal counter, and when the counter is zero the next coroutine waits until another one leaves. To limit concurrent API jobs, hold the semaphore for each job's whole life, from the submit to its final status, not just for the HTTP request.
The asyncio facts below come from the Python documentation, read 2026-09-28. The job API in the example is Sume's; its limits come from the generation admission docs.
How does asyncio.Semaphore work?
Python's docs: "A semaphore manages an internal counter which is decremented by each acquire() call and incremented by each release() call." The counter never goes below zero; when acquire() finds it at zero, it blocks until some task calls release(). The starting value defaults to 1, and "the preferred way to use a Semaphore is an async with statement", which releases even when the block raises (asyncio).
Prefer asyncio.BoundedSemaphore: it raises ValueError if release() would push the counter above its starting value, which turns a double release into an error instead of a silent extra slot. Neither class is thread-safe, so use them inside one event loop.
How do I limit concurrent API jobs with a semaphore?
Put the submit and the polling inside the same async with block. The example uses HTTPX's AsyncClient (HTTPX) against POST /v1/videos, which returns a job with a polling_url at once (Video generation). Each item gets its own stable Idempotency-Key, so rerunning the script with the same prompts replays accepted jobs instead of paying twice.
import asyncio, os, httpx
API = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
DONE = ("completed", "failed", "cancelled")
sem = asyncio.BoundedSemaphore(4) # your workspace's concurrency_limit
async def make_video(client, item_id, prompt):
async with sem: # held from submit to final status
r = await client.post(API, headers={**AUTH, "Idempotency-Key": f"batch-7-{item_id}"},
json={"model": "seedance-2", "prompt": prompt, "resolution": "720p"})
r.raise_for_status()
job = r.json()
while job["status"] not in DONE:
await asyncio.sleep(30)
job = (await client.get(job["polling_url"], headers=AUTH)).json()
return job
async def main(prompts):
async with httpx.AsyncClient() as client:
return await asyncio.gather(*(make_video(client, i, p) for i, p in enumerate(prompts)))Why hold the semaphore until the job finishes?
Because the POST returns in a moment, and the job runs for minutes. If the semaphore wraps only the request, a 500-prompt batch is submitted within seconds, and the semaphore limits nothing that matters.
Sume accepts jobs beyond your processing limit as queued, since "concurrency is a dispatch limit, not a submit limit". Every accepted job reserves its estimated cost at submit, and once the queue is also full, new submits fail with 429 queue_full (Generation admission). Holding the slot until the final status keeps in-flight work, and the balance reserved for it, at your processing limit.
What number should the semaphore be?
Start from your workspace's effective concurrency_limit; the docs call the dashboard Concurrency tab its source of truth (Generation admission). Video job concurrency and queueing shows how to size from a live generation_limits snapshot when other processes share the workspace.
A semaphore only counts the coroutines in its own event loop. Two scripts that each create BoundedSemaphore(4) can have eight jobs in flight against the same workspace, so split the limit between them.
Polling spends the read budget, which is separate from the write budget that submits spend, so a status loop cannot 429 your own submits.
What should each job do when a submit fails?
Decide per error code, inside the semaphore, so a failing job gives its slot back. A client-side timeout does not cancel a job, so never resubmit one just because your coroutine gave up.
| Response | Meaning | What the coroutine should do |
|---|---|---|
429 queue_full | No accepted-job capacity left | Stop adding work, wait for a job to finish, then submit under a new key; in current code a same-key resend returns the same queue_full |
429 rate_limited | Request budget exceeded | Wait for retry-after, then retry with the same key |
402 insufficient_credits | Balance can't cover the estimate | Stop the batch |
409 idempotency_conflict | Key reused with a different body | Fix the key scheme; reuse keys only for exact retries |
Sources
Related posts
More in Developers
- Bearer token vs API key: what's the difference?
An API key is a kind of credential; Bearer is a way to send one, in the Authorization header. A key can travel as a bearer token, as OAuth tokens do.
- Circuit breaker pattern in Python for AI API calls
A circuit breaker stops calling a failing API: after repeated failures it opens and fails fast, then lets a trial call through. A Python version.
- How to create an SRT file from text: time it with TTS
An SRT file needs a start and end time for every line. Voice your text with TTS word timings, then write each timed sentence as a numbered block.
- Creatify API: AI avatar videos, Aurora lip sync and credits
The Creatify API makes avatar videos from text or audio, Aurora talking videos from a photo and audio, and ads from a URL, billed in plan credits.
Written by Sume