503 provider_capacity_exceeded on a transcription submit: retry?
provider_capacity_exceeded means Sume's dispatch queue is full. Retry later with the same idempotency key. It differs from 429 queue_full and 402.

If a Sume transcription submit returns 503 provider_capacity_exceeded, Sume's provider dispatch queue is full. Retry later with the same Idempotency-Key, unless the error says not to retry. Nothing was accepted, so a retry with the same key is safe and cannot double-bill. This is different from 429 queue_full, which is your own workspace's accepted-job capacity. Both rules are in Errors and rate limits and Generation admission, read 2026-10-06.
Five refusals, five different reactions
A batch client should not treat every non-2xx the same. The table lists what each one means for a bulk transcription run.
| Status and code | Meaning | Client behavior |
|---|---|---|
503 provider_capacity_exceeded | Sume's provider dispatch queue is full | Retry later, same key |
429 queue_full | Your workspace has no queue room | Wait for a job to finish or cancel queued ones |
429 rate_limited | Request rate over the limit | Sleep for retry-after |
402 insufficient_credits | Balance cannot cover the estimate | Stop; add funds or shrink the request |
409 idempotency_conflict | Key reused with a different body | Stop; fix the key |
A classifier you can unit test
Keep the decision in one pure function so the loop around it stays small. The delays are your own choice, because the docs only say to retry later.
def classify(status, body, headers):
code = (body.get("error") or {}).get("code", "")
if status == 503 and code == "provider_capacity_exceeded":
return "retry_later_same_key", float(headers.get("retry-after", 30))
if status == 429 and code == "queue_full":
return "wait_for_a_job_to_finish", 15.0
if status == 429:
return "retry_after", float(headers.get("retry-after", 5))
if status == 402:
return "stop_add_funds", 0.0
if status == 409 and code == "idempotency_conflict":
return "stop_key_reused_with_new_body", 0.0
return "raise", 0.0
print(classify(503, {"error": {"code": "provider_capacity_exceeded"}}, {}))
print(classify(429, {"error": {"code": "queue_full"}}, {}))
print(classify(402, {"error": {"code": "insufficient_credits"}}, {}))How long to wait
Use exponential backoff with jitter, starting near 10 seconds, and send the same key every time. If you see retry-after, prefer it. Keep a cap on attempts per clip so a prolonged capacity problem pauses your batch instead of spinning.
Sources
Related posts
More in Developers
- A/B test two AI video models by API with a stable hash bucket
Split production video jobs between two Sume model ids by hashing a job key, so retries keep the same arm. Python code, a duration table and a 1,000-key check.
- arq worker that polls an AI video job with defer_by in Python
An arq task reads Sume's job status once and enqueues itself again with _defer_by from next_poll_after_seconds, giving asyncio polling without a sleep loop.
- asyncio Semaphore size for Sume image batches: accepted capacity
Size the semaphore to what Sume accepts, concurrency plus queue: Free 6, Pro 24, Startup 48, Scale 120. A fake-submit test proves the peak never exceeds it.
- asyncio TaskGroup cancels siblings: poll many Sume jobs safely
A TaskGroup cancels every other poller when one raises. For a batch of Sume video jobs that abandons waits, not jobs. Catch inside the task and return results.
Written by Sume