Image batch in Python: read ratelimit headers and retry-after

Sume can send ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after on image calls. Back off on 429 in Python, and treat queue_full separately.

4 min readSume
All posts

When you loop over many image requests, read the retry-after header on a 429 and wait that long; do not guess. Sume's errors page says public API responses can include ratelimit-limit, ratelimit-remaining, ratelimit-reset, and retry-after, tells you to back off on 429, and says to use retry-after when it is present. Two different 429 codes need different handling: rate_limited means too many requests in the window, and queue_full means the workspace has no room for another paid generation job.

The sketch below treats a numeric retry-after as seconds, which is how HTTP defines it, and falls back to exponential backoff with jitter when the header is absent.

Which 429 is it?

Read error.code from the body before you choose a delay. A rate limit clears with time. A full queue clears only when an existing queued or processing job finishes or is canceled, so shortening your sleep does not help.

The two 429 codes on submit. Source: docs.sume.com errors page and generation admission page, read 2026-10-03.
CodeMeaningWhat to do
rate_limitedToo many requests in the current windowBack off using retry-after when present
queue_fullWorkspace concurrency plus queue capacity is fullWait for jobs to finish or cancel queued jobs, then retry with the same idempotency key
402 insufficient_creditsBalance cannot reserve the estimated costNot a retry case; add funds or lower the request cost

How do I write the wait function?

Keep it pure so you can test it without a network call. The function takes the status, the headers, the error code, and the attempt number, and returns a number of seconds, or None when the response is not a rate-limit case.

import random

def wait_seconds(status, headers, error_code, attempt):
    if status != 429:
        return None
    h = {k.lower(): v for k, v in headers.items()}
    if error_code == "queue_full":
        return min(30 * (attempt + 1), 120)  # a job has to finish first
    try:
        return float(h["retry-after"])
    except (KeyError, ValueError):
        return min(2 ** attempt, 60) + random.random()

print(wait_seconds(429, {"Retry-After": "7"}, "rate_limited", 0))  # 7.0
print(wait_seconds(200, {}, None, 0))  # None

What must stay the same across retries?

The Idempotency-Key. The errors page says not to retry unsafe submit requests without one, and the same key on a retry returns the original job instead of billing a second one. Generate it once per image you want, not once per attempt.

Also remember that concurrency being full is not an error on its own. Sume accepts valid jobs as queued while queue capacity remains, so a batch can submit more jobs than your plan runs at once. Read generation_limits in the submit response for the current capacity, and prefer it over a static table.

How should the loop be structured?

Submit with mode: "async" or webhook for large batches, store each job id, and poll GET /v1/jobs/{id}/status with backoff. The Image API page says POST /v1/images blocks for up to 30 seconds and returns 202 with a job envelope when the work outlives that, so a long batch is better served by jobs from the start.

Poll reads can be rate-limited as well. The generation admission page calls those read and status limits polling backpressure, separate from generation concurrency. Spread polls out, use next_poll_after_seconds when the response gives it, and stop polling a job once it reports a terminal status.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume