Pydantic AI slot leak vs Sume queue_full: tell them apart
Pydantic AI v2.53.0 fixed a streamed-request concurrency slot leak. A client limiter is not Sume's workspace queue_full 429, and each needs its own handling.

Pydantic AI v2.53.0 patched a bug in ConcurrencyLimitedModel where streamed requests could keep their concurrency slot when released on a different task. That limiter lives in your process, while Sume's queue_full is a 429 from the API when your workspace has no generation capacity left; one is fixed by upgrading, the other by waiting for jobs to finish.
Two different limits
A client-side limiter caps how many calls your code starts at once. A server-side limit caps what the platform will accept. Mixing them up leads to the wrong fix, such as adding retries to a leaked slot or restarting the process because the queue is full.
| Aspect | Pydantic AI ConcurrencyLimitedModel | Sume queue_full |
|---|---|---|
| Where it lives | Your process | Sume API, per workspace |
| Symptom | Calls wait forever or stall locally | 429 with code queue_full |
| Cause named in v2.53.0 notes | Streamed request kept its slot when released on a different task | Concurrency plus queue capacity is full |
| Fix | Upgrade to a release with the patch | Wait for a job to finish or cancel queued jobs, retry with the same key |
What queue_full means on Sume
Sume's errors page says queue_full is different from ordinary request rate limiting: Sume cannot accept another paid generation job for the workspace until an existing queued or processing job finishes or is canceled. Concurrency being full is not an error by itself, since valid jobs are accepted as queued while queue capacity remains.
rate_limited is a separate 429 about request volume; back off with retry-after when present. Do not retry an unsafe submit without an Idempotency-Key.
Handle each in its own place
Keep your limiter small and release it in a finally block on the same task that acquired it. For Sume, retry on queue_full with the same Idempotency-Key so the retry returns the original job if the first submit was in fact accepted. The sketch needs pip install httpx.
import asyncio, os
import httpx
limit = asyncio.Semaphore(4) # your client-side cap
async def submit(client, key, body):
async with limit:
for attempt in range(5):
r = await client.post("/v1/videos", json=body,
headers={"Idempotency-Key": key})
if r.status_code != 429:
return r.json()
code = r.json().get("error", {}).get("code")
wait = float(r.headers.get("retry-after", 2 ** attempt))
print("429", code, "waiting", wait)
await asyncio.sleep(wait)
raise RuntimeError("still 429 after retries")
async def main():
headers = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
async with httpx.AsyncClient(base_url="https://api.sume.com", headers=headers) as client:
body = {"model": "sume/auto", "prompt": "A mug on a desk, soft light"}
print(await submit(client, "demo-001", body))
asyncio.run(main())Sources
Related posts
More in Developers
- Pydantic AI Workspace sandboxes: fetch Sume artifacts
Pydantic AI's Workspace abstraction runs tools locally or in sandboxes. Inside a sandbox, fetch Sume artifact URLs, and give Sume inputs as public HTTPS URLs.
- Python 3.10 is end of life: a stdlib Sume webhook verifier
Python 3.10 has reached end of life. A standard-library verifier for Sume's signed webhooks that refuses an empty secret and accepts rotated signatures.
- Python 3.14.8 urllib credential fix: a stdlib Sume job poll script
Python 3.14.8 fixed urllib HTTPPasswordMgr handing credentials across schemes. Upgrade, and call Sume with an x-api-key header from a plain stdlib script.
- Railway webhook custom headers vs Sume's signed headers
Railway project webhooks now accept custom headers. Sume signs its callbacks with x-sume-webhook-timestamp and x-sume-webhook-signature; verify them in Python.
Written by Sume