rate_limited or queue_full? One Python submit handler for both 429s

Two different 429s need two different waits. A Python handler reads error.code, sleeps on retry-after for rate_limited, and waits for capacity on queue_full.

3 min readSume
All posts

Both errors are HTTP 429, so branch on error.code, not the status. rate_limited means you sent too many requests in the current window: sleep for retry-after and resend. queue_full means the workspace has no accepted generation capacity left: sleeping a second will not help, because a running job has to finish or a queued one has to be canceled first.

In both cases resend with the same Idempotency-Key. The docs say not to retry unsafe submits without one, and with one the retry cannot create a second paid job.

The handler

Any other error code raises at once with the request_id, which is the value Sume support asks for.

import os, time, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def submit(body, key, tries=6):
    for n in range(tries):
        r = requests.post("https://api.sume.com/v1/videos", json=body,
                          headers={**H, "Idempotency-Key": key}, timeout=30)
        if r.ok:
            return r.json()
        err = r.json().get("error", {})
        code = err.get("code")
        if code == "rate_limited":
            time.sleep(float(r.headers.get("retry-after", 2 ** n)))
        elif code == "queue_full":
            time.sleep(min(60, 15 * (n + 1)))  # let a running job finish
        else:
            raise RuntimeError("%s request_id=%s" % (code, err.get("request_id")))
    raise RuntimeError("still limited after %d tries; key %s" % (tries, key))

The two 429s side by side

Sume docs, read 2026-10-08
CodeCauseRight move
rate_limitedRequest volume over an abuse-protection limitBack off for retry-after, then resend
queue_fullProcessing seats and queue slots are all usedWait for a terminal job or cancel queued ones, then resend
402 insufficient_creditsThe reserve does not fit the balanceDo not retry; send a cheaper request, wait for included Gen$ or upgrade the plan

Do not confuse it with polling limits

Status and list reads have their own limits. A 429 on a poll is read backpressure, not generation concurrency, and the fix is a longer poll interval, not fewer jobs.

Waiting for capacity properly

The fixed sleep for queue_full in the sample is deliberately simple. A better version reads the status of the jobs you already submitted and resumes as soon as one is terminal. That reacts to real capacity instead of guessing, and it never adds load while the queue is full.

  • Keep a set of in-flight job ids in your process.
  • On queue_full, poll those ids with backoff until one reports terminal.
  • Then resend the rejected submit with the same key.
  • If you no longer need some queued work, cancel it; a queued job cancels cleanly and releases its reserve.

Use the headers

Headers help with the other 429. Public responses can carry ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after; when retry-after is present, use it in preference to your own schedule.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume