rate_limited or queue_full? One Python submit handler for both 429s
Two different 429s need two different waits. A Python handler reads error.code, sleeps on retry-after for rate_limited, and waits for capacity on queue_full.

Both errors are HTTP 429, so branch on error.code, not the status. rate_limited means you sent too many requests in the current window: sleep for retry-after and resend. queue_full means the workspace has no accepted generation capacity left: sleeping a second will not help, because a running job has to finish or a queued one has to be canceled first.
In both cases resend with the same Idempotency-Key. The docs say not to retry unsafe submits without one, and with one the retry cannot create a second paid job.
The handler
Any other error code raises at once with the request_id, which is the value Sume support asks for.
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def submit(body, key, tries=6):
for n in range(tries):
r = requests.post("https://api.sume.com/v1/videos", json=body,
headers={**H, "Idempotency-Key": key}, timeout=30)
if r.ok:
return r.json()
err = r.json().get("error", {})
code = err.get("code")
if code == "rate_limited":
time.sleep(float(r.headers.get("retry-after", 2 ** n)))
elif code == "queue_full":
time.sleep(min(60, 15 * (n + 1))) # let a running job finish
else:
raise RuntimeError("%s request_id=%s" % (code, err.get("request_id")))
raise RuntimeError("still limited after %d tries; key %s" % (tries, key))The two 429s side by side
| Code | Cause | Right move |
|---|---|---|
rate_limited | Request volume over an abuse-protection limit | Back off for retry-after, then resend |
queue_full | Processing seats and queue slots are all used | Wait for a terminal job or cancel queued ones, then resend |
402 insufficient_credits | The reserve does not fit the balance | Do not retry; send a cheaper request, wait for included Gen$ or upgrade the plan |
Do not confuse it with polling limits
Status and list reads have their own limits. A 429 on a poll is read backpressure, not generation concurrency, and the fix is a longer poll interval, not fewer jobs.
Waiting for capacity properly
The fixed sleep for queue_full in the sample is deliberately simple. A better version reads the status of the jobs you already submitted and resumes as soon as one is terminal. That reacts to real capacity instead of guessing, and it never adds load while the queue is full.
- Keep a set of in-flight job ids in your process.
- On
queue_full, poll those ids with backoff until one reportsterminal. - Then resend the rejected submit with the same key.
- If you no longer need some queued work, cancel it; a queued job cancels cleanly and releases its reserve.
Use the headers
Headers help with the other 429. Public responses can carry ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after; when retry-after is present, use it in preference to your own schedule.
Sources
Related posts
More in Developers
- Test one video prompt on ten models for under $7 (Python sweep)
A 5-second, lowest-resolution sweep of one prompt over ten Sume video ids costs about $5.60 in total. The price of each row and a script that submits them.
- TTS word timestamps: timestamps.words and sentence segmentation
Sume TTS accepts timestamps.words and segmentation.mode sentence so a generated voiceover can drive caption timing. Request fields, rules and a working call.
- Turn a roleplay debrief into an avatar feedback clip in Python
Take the written debrief from a roleplay or survey session and render it as a 16:9 Sume avatar clip with a retry-safe key, a 12 to 168 word check and polling.
- 12 Wan 3.0 clips in parallel in Python: ThreadPoolExecutor, width 4
A Python batch for Sume: ThreadPoolExecutor at width 4 (Pro concurrency), one Idempotency-Key per item, polling by next_poll_after_seconds. Cost included.
Written by Sume