503 provider_capacity_exceeded on a transcription submit: retry?

provider_capacity_exceeded means Sume's dispatch queue is full. Retry later with the same idempotency key. It differs from 429 queue_full and 402.

4 min readSume
All posts

If a Sume transcription submit returns 503 provider_capacity_exceeded, Sume's provider dispatch queue is full. Retry later with the same Idempotency-Key, unless the error says not to retry. Nothing was accepted, so a retry with the same key is safe and cannot double-bill. This is different from 429 queue_full, which is your own workspace's accepted-job capacity. Both rules are in Errors and rate limits and Generation admission, read 2026-10-06.

Five refusals, five different reactions

A batch client should not treat every non-2xx the same. The table lists what each one means for a bulk transcription run.

Submit refusals and client behavior, from Errors and rate limits and Generation admission (docs.sume.com), read 2026-10-06.
Status and codeMeaningClient behavior
503 provider_capacity_exceededSume's provider dispatch queue is fullRetry later, same key
429 queue_fullYour workspace has no queue roomWait for a job to finish or cancel queued ones
429 rate_limitedRequest rate over the limitSleep for retry-after
402 insufficient_creditsBalance cannot cover the estimateStop; add funds or shrink the request
409 idempotency_conflictKey reused with a different bodyStop; fix the key

A classifier you can unit test

Keep the decision in one pure function so the loop around it stays small. The delays are your own choice, because the docs only say to retry later.

def classify(status, body, headers):
    code = (body.get("error") or {}).get("code", "")
    if status == 503 and code == "provider_capacity_exceeded":
        return "retry_later_same_key", float(headers.get("retry-after", 30))
    if status == 429 and code == "queue_full":
        return "wait_for_a_job_to_finish", 15.0
    if status == 429:
        return "retry_after", float(headers.get("retry-after", 5))
    if status == 402:
        return "stop_add_funds", 0.0
    if status == 409 and code == "idempotency_conflict":
        return "stop_key_reused_with_new_body", 0.0
    return "raise", 0.0


print(classify(503, {"error": {"code": "provider_capacity_exceeded"}}, {}))
print(classify(429, {"error": {"code": "queue_full"}}, {}))
print(classify(402, {"error": {"code": "insufficient_credits"}}, {}))

How long to wait

Use exponential backoff with jitter, starting near 10 seconds, and send the same key every time. If you see retry-after, prefer it. Keep a cap on attempts per clip so a prolonged capacity problem pauses your batch instead of spinning.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume