Run one prompt on five image models via Sume: 200 vs 202 in Python
A Python script fires one prompt at Seedream, FLUX.2, Qwen, Ideogram and Recraft ids on Sume and handles both the 200 image reply and the 202 job envelope.

To run one prompt across Seedream, FLUX.2, Qwen-Image, Ideogram and Recraft on Sume, post the same body to POST /v1/images with a different model each time. A reply is either 200 with the images inline or 202 with a job envelope you poll, so the script must branch on the status code, not on the body shape.
Every behaviour below is from the Image API and Jobs and results pages.
Why do some replies come back as 202?
POST /v1/images blocks for up to 30 seconds and returns 200 with the image response when the generation finishes in time. If it is still running at the budget, or you sent mode: "async", or mode: "webhook" with a webhook_url, Sume returns 202 with a job envelope holding a status_url and a result_url. The page's own advice is to check the status code, not the body shape, and notes that 4K, high quality and large n are the most likely to degrade to 202.
What does the script do?
It runs the calls in worker threads, since requests is blocking, collects each reply, and for any 202 polls GET /v1/jobs/{id}/status until terminal, then reads GET /v1/jobs/{id}/result. The result shape is the standard job result with result.artifacts[], not the inline image body.
import asyncio, os, time, requests
B = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
MODELS = ["bytedance-seed/seedream-4.5", "black-forest-labs/flux.2-pro",
"qwen/qwen-image", "ideogram/ideogram-v3", "recraft/recraft-v4"]
def run(m):
r = requests.post(B + "/v1/images", headers=H, timeout=60,
json={"model": m, "prompt": "A red bicycle in a rainy alley"})
if r.status_code == 200:
return m, r.json()["data"][0]["url"]
if r.status_code != 202:
return m, f"HTTP {r.status_code}"
jid = r.json()["data"]["job"]["id"]
while True:
s = requests.get(f"{B}/v1/jobs/{jid}/status", headers=H, timeout=30).json()
if str(s).find('"completed"') >= 0 or str(s).find('"failed"') >= 0:
return m, requests.get(f"{B}/v1/jobs/{jid}/result", headers=H, timeout=30).json()
time.sleep(3)
async def main():
for m, out in await asyncio.gather(*(asyncio.to_thread(run, x) for x in MODELS)):
print(m, out)
asyncio.run(main())What can go wrong when you fan out?
Five concurrent submissions count against your workspace. The error table separates 429 rate_limited, a request-window limit, from 429 queue_full, where the workspace's concurrency plus queue is full. Back off on both and use retry-after when it is present. Do not retry unsafe submits without an Idempotency-Key; add a distinct one per model call if you wrap this in a retry loop.
A 402 insufficient_credits is not a retry case at all: stop and top up. Failed generations are not billed, so a run that ends in errors costs only what completed.
| Status | Meaning for this script | Action |
|---|---|---|
| 200 | Images inline in data[] | Read data[0].url and usage.cost |
| 202 | Job envelope | Poll status_url, then result_url |
| 402 | insufficient_credits | Stop, top up, resume |
| 429 | rate_limited or queue_full | Back off, honor retry-after |
| 400 | unsupported_parameter or invalid request | Fix the body; check the model descriptor |
Which ids should you use?
The ids above are the slugs in the repository's image contract. Your live catalog is authoritative, so list it first or diff it on a schedule and edit the MODELS list when something is added or retired. For a fuller submit-and-poll walkthrough with httpx see the asyncio post.
Extending the script
Cost per model comes back in usage.cost on 200 replies. On 202 replies read the job result instead; the two shapes differ, which is why the script branches.
- Switch to
mode: "webhook"with awebhook_urlwhen you do not want to hold threads. - Write each result URL and cost into a CSV so you can compare price against quality later.
- Add a per-model timeout and a maximum poll count so a stuck job cannot hang the run.
- Cancel jobs you no longer need; the status response includes a cancel URL when cancellation is possible.
Sources
Related posts
More in Developers
- Why concurrency_limit differs from the Sume plan table
In Sume generation_limits, concurrency_limit is the effective cap and limit_source says plan or admin_override. Size work from it, not plan_concurrency_limit.
- curl --retry on a POST: retry a Sume submit with one key
curl --retry also retries a POST, and it resends the same headers each time. Put an Idempotency-Key on a Sume submit first, then pick --retry-max-time.
- Decart lucy-latest vs a pinned Lucy model; Sume catalog ids
Decart's lucy-latest alias can move while legacy Lucy Clip costs $0.15 per second against $0.04 for Lucy 2.5. Why pin a model id, and how to do it on Sume.
- Dub a 20-minute video: audio detach 900 s cap, STT 600 s, TTS 1,200 s
A 20-minute video needs chunking before a dub: audio detach outputs at most 900 s, STT reservation tops out at 600 s, and TTS fails past 1,200 s.
Written by Sume