Fan out 40 image variants without hitting queue_full

Sume accepts jobs as queued up to your plan's capacity, then returns 429 queue_full. Here are the plan numbers and a wave loop that reads generation_limits.

4 min readSume
All posts

To submit 40 image variants, check your workspace's accepted-job capacity first. Sume queues valid paid jobs as queued while capacity remains, and returns 429 queue_full when the workspace has no accepted capacity left. On the Free plan that capacity is 6 jobs, on Pro 24, and on Startup 48, so 40 variants need Startup or a wave loop.

Concurrency is a dispatch limit, not a submit limit: jobs wait in queued and move to processing as slots open. queued is not a failure.

What are the plan limits?

Queue capacity defaults to max(3, concurrency_limit x 5). The dashboard Concurrency tab and the generation_limits field are the source of truth; admin overrides can raise them.

Plan defaults from Sume's Generation admission docs, read 2026-10-02
PlanProcessingQueueAccepted
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

How do I send waves?

Submit responses can carry a generation_limits snapshot with queue_capacity_remaining and a wave_size_hint, defined as max(1, floor(queue_capacity_remaining x 0.75)). The hint sizes a submission wave only; do not use it as a concurrency limit. The sketch below stops a wave on queue_full and uses a different idempotency key per variant.

import os
import requests

URL = "https://api.sume.com/v1/images"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
variants = [f"Product on a {c} backdrop" for c in ["red", "blue", "sage"]]
pending = []

for i, prompt in enumerate(variants):
    body = {"model": "sume/auto", "prompt": prompt, "mode": "async"}
    headers = {**HEADERS, "Idempotency-Key": f"variants-001-{i}"}
    r = requests.post(URL, headers=headers, json=body, timeout=60)
    if r.status_code == 429:
        print("stop wave:", r.json()["error"]["code"])
        break
    r.raise_for_status()
    limits = r.json().get("generation_limits") or {}
    print(i, r.status_code, limits.get("queue_capacity_remaining"))
    pending.append(i)

What do I do after queue_full?

Wait for jobs to finish or cancel queued ones, then retry with the same idempotency key. Do not confuse it with 429 rate_limited, which is request volume and calls for backoff using retry-after. Poll status with backoff and fetch results only when result_ready is true. See Generation admission.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume