Fan out 40 image variants without hitting queue_full
Sume accepts jobs as queued up to your plan's capacity, then returns 429 queue_full. Here are the plan numbers and a wave loop that reads generation_limits.

To submit 40 image variants, check your workspace's accepted-job capacity first. Sume queues valid paid jobs as queued while capacity remains, and returns 429 queue_full when the workspace has no accepted capacity left. On the Free plan that capacity is 6 jobs, on Pro 24, and on Startup 48, so 40 variants need Startup or a wave loop.
Concurrency is a dispatch limit, not a submit limit: jobs wait in queued and move to processing as slots open. queued is not a failure.
What are the plan limits?
Queue capacity defaults to max(3, concurrency_limit x 5). The dashboard Concurrency tab and the generation_limits field are the source of truth; admin overrides can raise them.
| Plan | Processing | Queue | Accepted |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
How do I send waves?
Submit responses can carry a generation_limits snapshot with queue_capacity_remaining and a wave_size_hint, defined as max(1, floor(queue_capacity_remaining x 0.75)). The hint sizes a submission wave only; do not use it as a concurrency limit. The sketch below stops a wave on queue_full and uses a different idempotency key per variant.
import os
import requests
URL = "https://api.sume.com/v1/images"
HEADERS = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
variants = [f"Product on a {c} backdrop" for c in ["red", "blue", "sage"]]
pending = []
for i, prompt in enumerate(variants):
body = {"model": "sume/auto", "prompt": prompt, "mode": "async"}
headers = {**HEADERS, "Idempotency-Key": f"variants-001-{i}"}
r = requests.post(URL, headers=headers, json=body, timeout=60)
if r.status_code == 429:
print("stop wave:", r.json()["error"]["code"])
break
r.raise_for_status()
limits = r.json().get("generation_limits") or {}
print(i, r.status_code, limits.get("queue_capacity_remaining"))
pending.append(i)What do I do after queue_full?
Wait for jobs to finish or cancel queued ones, then retry with the same idempotency key. Do not confuse it with 429 rate_limited, which is request volume and calls for backoff using retry-after. Poll status with backoff and fetch results only when result_ready is true. See Generation admission.
Sources
Related posts
More in Developers
- Find near-duplicate generated ad images with an average hash in Python
Four images from one prompt can be almost the same picture. A 64-bit average hash in Pillow flags the near-duplicates so you review each look once.
- Fit narration to a 30-second slot: tune TTS speed from measured length
Measure a TTS take with timestamps.words, compute the speed that hits 30 seconds, and know when to cut words instead. Python, within Sume's 0.6 to 1.5 range.
- Fix one sentence in an AI avatar video without a full re-render
ElevenLabs can regenerate only edited dubbing regions. A Sume avatar video is one job, so the fix is to keep clips short and join them: the cost math.
- FLUX 3 pixel-exact local edits vs the Sume mask_url edit
FLUX 3 Image edits several marked elements in one request. Sume's mask_url edit is documented for ChatGPT Image 2.5 only; here is how to run a local edit.
Written by Sume