AI image API 429s: queue_full vs rate_limited, and how to retry each
Sume returns 429 for two different reasons. rate_limited means back off; queue_full means wait for jobs to finish. A Python retry that treats them differently.

Both errors are HTTP 429, but you handle them differently. rate_limited means request volume passed an abuse-protection limit: back off, using retry-after when it is present. queue_full means the workspace has no accepted-job capacity left: wait for jobs to finish or cancel queued ones, then retry with the same idempotency key (generation admission docs). Retrying queue_full in a tight loop only adds noise.
The reason to separate them is bulk image runs. A 200-prompt script on a Free plan will hit queue_full quickly because only 6 jobs can be accepted at once, while a fast loop on any plan can trip rate_limited.
How much capacity do I have?
Defaults from the docs, read 2026-10-01.
| Plan | Processing | Queue | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
What does the retry look like?
The same Idempotency-Key is sent on every attempt for one prompt, so a retried submit cannot create a second paid job.
import os
import time
import uuid
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def submit(prompt: str, tries: int = 6):
key = str(uuid.uuid4())
for n in range(tries):
r = requests.post("https://api.sume.com/v1/images", timeout=60,
headers={**H, "Idempotency-Key": key},
json={"model": "openai/gpt-image-2.5", "prompt": prompt})
if r.status_code != 429:
return r
wait = float(r.headers.get("retry-after", 2 ** n))
code = r.json().get("error", {}).get("code")
print(code, "waiting", wait)
time.sleep(wait * (5 if code == "queue_full" else 1))
raise RuntimeError("still 429")What do I do when the queue is the problem?
Slow the producer. Send no more than the processing plus queue number at a time, or use the async job flow and poll next_poll_after_seconds rather than holding sync requests open. The sync call waits up to 30 seconds and then returns a 202 envelope, so long waits do not need long timeouts.
Why the idempotency key matters
A submit that times out on your side may still have been accepted. If you retry without a key, you can pay for two images. With the same Idempotency-Key on each attempt, the retry maps to the same job. Generate one key per prompt, not per attempt, and store it with the prompt so a restarted script can resume.
Limits
The 5x multiplier on queue_full is a choice in this sketch, not a documented number; a better rule is to check job status and resubmit when capacity frees. Enterprise or admin overrides may change the limits above, so read your own concurrency_limit from the account. The error body shape here follows the docs example, but print r.text once to confirm it on your key.
Sources
Related posts
More in Developers
- Is there an asset library API for AI images and videos?
Sume has no folders or tags. Your library is completed jobs plus durable media.sume.com artifacts, which you list, label by Idempotency-Key, and download.
- Claude Cost Report API: daily buckets by workspace vs Sume /v1/usage
Anthropic's cost_report endpoint returns USD cost in 1d buckets, groupable by workspace or description. Sume's /v1/usage sums one thread, run or job instead.
- Avatar catalog search: explore mode, seed and diversity
POST /v1/avatar-catalog/search browses reusable avatars. Omit the query for explore mode, pass a seed to keep the order stable, and set diversity from 0 to 1.
- Avatar catalog search returns few results: auto_expand explained
When an avatar catalog search is thin, Sume relaxes filters in a set order and lists them in relaxed_filters. Set auto_expand to false for strict matching.
Written by Sume