Image burst after a model launch: 429 queue_full vs 503 retry plan

Launch week means batches. On Sume, 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded each need a different retry, plus idempotency keys.

4 min readSume
All posts

When a new image model launches, you queue a few hundred prompts at once. Sume answers with different errors that need different retries: 429 rate_limited means back off using retry-after, 429 queue_full means wait for a running job to finish, and 503 provider_capacity_exceeded means retry later with the same idempotency key. Never re-submit a paid request under a new key.

The codes that matter

Sume's errors page lists these codes. The point of separating them is that two look alike (both are 429) but need opposite handling: a rate limit is about request frequency, and queue_full is about your workspace's capacity for generation jobs.

Burst-time errors and the retry each needs (Sume docs, read 2026-10-07) (read 2026-10-07)
Status and codeMeaningRetry
429 rate_limitedToo many requests in the windowBack off; honor retry-after if present
429 queue_fullConcurrency and queue capacity are both fullWait until a queued or processing job ends or is canceled
503 provider_capacity_exceededSume's provider dispatch queue is fullRetry later with the same idempotency key
402 insufficient_creditsBalance too low for the generationTop up; do not retry in a loop
400 invalid_requestBad body or parameterFix it; a retry will fail the same way

One key per intent

The Jobs page says to retry a submit with the same Idempotency-Key so the retry returns the original job instead of billing a second one. Build the key from the prompt row, for example the SKU and the shot name, so a crash and restart reuses it. Use the key again only for the same operation and payload.

A retry loop

The loop below uses mode: "async" so each submit returns a job quickly. It retries only on 429 and 503 and gives up on anything else.

import os
import time
import requests

URL = "https://api.sume.com/v1/images"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}


def submit(key, body):
    headers = dict(HEAD, **{"Idempotency-Key": key})
    delay = 2
    for _ in range(6):
        r = requests.post(URL, headers=headers, json=body, timeout=45)
        if r.status_code in (200, 202):
            return r.json()
        if r.status_code not in (429, 503):
            raise RuntimeError("status %s: %s" % (r.status_code, r.text[:200]))
        wait = float(r.headers.get("retry-after", delay))
        time.sleep(wait)
        delay = min(delay * 2, 60)
    raise RuntimeError("gave up on " + key)


def main():
    body = {"model": "ideogram/ideogram-v4.5", "prompt": "launch banner", "mode": "async"}
    print(submit("banner-sku-001-v1", body))


main()

Keep the batch small enough

queue_full is a signal to submit fewer jobs at once, not to retry harder. Cap in-flight jobs on your side, poll the status URLs, and add the next prompt as a job finishes. The Generation admission page explains how queued state and concurrency interact.

What this does not cover

A failed generation is not billed, but a 402 is a balance problem, and retries will not fix it. I have not measured how long launch-week queues last, so the delays above are defaults you should tune.

Practical settings for a launch-week run:

  • Cap in-flight jobs on your side.
  • Log the request id from every error body.
  • Alert on repeated 402, because retries will not clear it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume