Launch-week 503 provider_capacity_exceeded: safe video submit retries

New video models cause capacity spikes. Retry 429 and 503 on Sume with the same Idempotency-Key, honor retry-after, never retry 402. Python sample included.

5 min readSume
All posts

When a new video model launches and traffic spikes, Sume can answer a submit with 429 or 503. Retry both with the same Idempotency-Key, wait for retry-after when it is present, and never retry a 402, which means the balance is short.

The script below implements that in under 30 lines. The rules come from Sume's Errors and rate limits page.

What each status means

The Errors and rate limits page lists 429 rate_limited for too many requests in the window, 429 queue_full when workspace concurrency and queue capacity are both full, and 503 provider_capacity_exceeded when Sume's provider dispatch queue is full. For the last it says to retry later with the same idempotency key.

A 402 insufficient_credits is different: the balance does not cover the reservation, and a retry only repeats the failure.

Submit errors and what to do (Sume docs, checked 2026-10-10)
Status and codeMeaningRetry?
429 rate_limitedToo many requests in the windowYes, with backoff
429 queue_fullWorkspace queue is fullYes, after a job finishes
503 provider_capacity_exceededProvider dispatch queue is fullYes, same key
503 provider_not_configuredProvider unavailable in this runtimeNo, check the catalog
402 insufficient_creditsBalance too lowNo, add funds

Why the key must be reused

The docs warn against retrying unsafe submits without an Idempotency-Key. A replay with the same key returns the original job, so if your first request reached Sume but the response was lost, the retry does not create a second job or a second reservation. A fresh key per attempt is the mistake that doubles a bill.

Generate the key once per logical job, before the first attempt, and reuse it for every retry of that job.

The script

The loop retries only 429 and 503. It sleeps for retry-after if the header is present, otherwise it backs off exponentially up to two minutes. Any other error status raises immediately, so a 402 or a 400 stops the loop. Sleep times are examples; tune them to your volume.

import os, time, uuid, requests

API = "https://api.sume.com/v1/videos"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def submit(body, tries=6):
    key = str(uuid.uuid4())  # one key for every retry of this job
    for n in range(tries):
        r = requests.post(API, json=body, timeout=60,
                          headers={**H, "Idempotency-Key": key})
        if r.status_code not in (429, 503):
            r.raise_for_status()  # 402 and 400 stop here
            return r.json()
        wait = float(r.headers.get("retry-after") or min(5 * 2 ** n, 120))
        print(r.status_code, "busy, retrying in", wait, "s")
        time.sleep(wait)
    raise RuntimeError("still busy after retries")

job = submit({"model": "wan-3.0", "prompt": "waves at dusk", "duration": 5})
print(job["status"], job["polling_url"])

Beyond retries

If a launch drives sustained load, retries only smooth short spikes. Keep your own concurrency below what the workspace queue can hold, and submit in small waves instead of a thousand at once. The error page notes that full concurrency alone is not an error while queue capacity remains, so Sume accepts valid jobs as queued.

Also consider spreading across rows. A model that just launched is often busy, while an older row may have headroom. Check the catalog for a row with the same inputs, and keep the model id in config so you can switch without a deploy.

What to log

Log the status, the Sume error code, the request_id from the error body and the attempt number for every retry. The errors page says the request id is safe to share with support, and it is the thing to quote if retries keep failing. Do not log the API key or signed URLs.

After a run, count how many jobs needed a retry and in which hour. If a pattern appears around a launch, you can schedule heavy batches away from it. A launch-day spike usually fades within days, so favour waiting over rewriting code.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume