Launch-week 503 provider_capacity_exceeded: safe video submit retries
New video models cause capacity spikes. Retry 429 and 503 on Sume with the same Idempotency-Key, honor retry-after, never retry 402. Python sample included.

When a new video model launches and traffic spikes, Sume can answer a submit with 429 or 503. Retry both with the same Idempotency-Key, wait for retry-after when it is present, and never retry a 402, which means the balance is short.
The script below implements that in under 30 lines. The rules come from Sume's Errors and rate limits page.
What each status means
The Errors and rate limits page lists 429 rate_limited for too many requests in the window, 429 queue_full when workspace concurrency and queue capacity are both full, and 503 provider_capacity_exceeded when Sume's provider dispatch queue is full. For the last it says to retry later with the same idempotency key.
A 402 insufficient_credits is different: the balance does not cover the reservation, and a retry only repeats the failure.
| Status and code | Meaning | Retry? |
|---|---|---|
| 429 rate_limited | Too many requests in the window | Yes, with backoff |
| 429 queue_full | Workspace queue is full | Yes, after a job finishes |
| 503 provider_capacity_exceeded | Provider dispatch queue is full | Yes, same key |
| 503 provider_not_configured | Provider unavailable in this runtime | No, check the catalog |
| 402 insufficient_credits | Balance too low | No, add funds |
Why the key must be reused
The docs warn against retrying unsafe submits without an Idempotency-Key. A replay with the same key returns the original job, so if your first request reached Sume but the response was lost, the retry does not create a second job or a second reservation. A fresh key per attempt is the mistake that doubles a bill.
Generate the key once per logical job, before the first attempt, and reuse it for every retry of that job.
The script
The loop retries only 429 and 503. It sleeps for retry-after if the header is present, otherwise it backs off exponentially up to two minutes. Any other error status raises immediately, so a 402 or a 400 stops the loop. Sleep times are examples; tune them to your volume.
import os, time, uuid, requests
API = "https://api.sume.com/v1/videos"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def submit(body, tries=6):
key = str(uuid.uuid4()) # one key for every retry of this job
for n in range(tries):
r = requests.post(API, json=body, timeout=60,
headers={**H, "Idempotency-Key": key})
if r.status_code not in (429, 503):
r.raise_for_status() # 402 and 400 stop here
return r.json()
wait = float(r.headers.get("retry-after") or min(5 * 2 ** n, 120))
print(r.status_code, "busy, retrying in", wait, "s")
time.sleep(wait)
raise RuntimeError("still busy after retries")
job = submit({"model": "wan-3.0", "prompt": "waves at dusk", "duration": 5})
print(job["status"], job["polling_url"])Beyond retries
If a launch drives sustained load, retries only smooth short spikes. Keep your own concurrency below what the workspace queue can hold, and submit in small waves instead of a thousand at once. The error page notes that full concurrency alone is not an error while queue capacity remains, so Sume accepts valid jobs as queued.
Also consider spreading across rows. A model that just launched is often busy, while an older row may have headroom. Check the catalog for a row with the same inputs, and keep the model id in config so you can switch without a deploy.
What to log
Log the status, the Sume error code, the request_id from the error body and the attempt number for every retry. The errors page says the request id is safe to share with support, and it is the thing to quote if retries keep failing. Do not log the API key or signed URLs.
After a run, count how many jobs needed a retry and in which hour. If a pattern appears around a launch, you can schedule heavy batches away from it. A launch-day spike usually fades within days, so favour waiting over rewriting code.
Sources
Related posts
More in Developers
- Let a browser poll your backend, not the Sume API: a proxy pattern
Keep SUME_API_KEY on the server: submit async, return the job id, and give the browser a read-only status route. A 27-line TypeScript route with the checks.
- Lint a Format package in CI before you PUT it
A short Python check for a Sume Format package: SKILL.md name equals the slug, allowed folders and extensions, file names, 100 MiB limits. Run it before PUT.
- Lip-sync a 2-minute monologue: split one TTS wav into H3 Max clips
MiniMax H3 Max lip-sync takes audio of 5 to 14.8 seconds. For a longer speech, render one TTS wav and cut it at sentence ends; the cost is worked below.
- Is there a lipsync-1.0 endpoint on Sume? Old paths 404; use Fabric
Sume's old /v1/lipsync-1.0 paths return 404 and its model ids return model_not_found. Send the same still and audio to veed/fabric-1.0 or H3 Max lip-sync.
Written by Sume