Pydantic AI slot leak vs Sume queue_full: tell them apart

Pydantic AI v2.53.0 fixed a streamed-request concurrency slot leak. A client limiter is not Sume's workspace queue_full 429, and each needs its own handling.

4 min readSume
All posts

Pydantic AI v2.53.0 patched a bug in ConcurrencyLimitedModel where streamed requests could keep their concurrency slot when released on a different task. That limiter lives in your process, while Sume's queue_full is a 429 from the API when your workspace has no generation capacity left; one is fixed by upgrading, the other by waiting for jobs to finish.

Two different limits

A client-side limiter caps how many calls your code starts at once. A server-side limit caps what the platform will accept. Mixing them up leads to the wrong fix, such as adding retries to a leaked slot or restarting the process because the queue is full.

Client limiter versus Sume workspace queue (read 2026-10-03)
AspectPydantic AI ConcurrencyLimitedModelSume queue_full
Where it livesYour processSume API, per workspace
SymptomCalls wait forever or stall locally429 with code queue_full
Cause named in v2.53.0 notesStreamed request kept its slot when released on a different taskConcurrency plus queue capacity is full
FixUpgrade to a release with the patchWait for a job to finish or cancel queued jobs, retry with the same key

What queue_full means on Sume

Sume's errors page says queue_full is different from ordinary request rate limiting: Sume cannot accept another paid generation job for the workspace until an existing queued or processing job finishes or is canceled. Concurrency being full is not an error by itself, since valid jobs are accepted as queued while queue capacity remains.

rate_limited is a separate 429 about request volume; back off with retry-after when present. Do not retry an unsafe submit without an Idempotency-Key.

Handle each in its own place

Keep your limiter small and release it in a finally block on the same task that acquired it. For Sume, retry on queue_full with the same Idempotency-Key so the retry returns the original job if the first submit was in fact accepted. The sketch needs pip install httpx.

import asyncio, os
import httpx

limit = asyncio.Semaphore(4)  # your client-side cap

async def submit(client, key, body):
    async with limit:
        for attempt in range(5):
            r = await client.post("/v1/videos", json=body,
                                  headers={"Idempotency-Key": key})
            if r.status_code != 429:
                return r.json()
            code = r.json().get("error", {}).get("code")
            wait = float(r.headers.get("retry-after", 2 ** attempt))
            print("429", code, "waiting", wait)
            await asyncio.sleep(wait)
        raise RuntimeError("still 429 after retries")

async def main():
    headers = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
    async with httpx.AsyncClient(base_url="https://api.sume.com", headers=headers) as client:
        body = {"model": "sume/auto", "prompt": "A mug on a desk, soft light"}
        print(await submit(client, "demo-001", body))

asyncio.run(main())

Sources

Related posts

More in Developers

All Developers posts

Written by Sume