30 image jobs in one async batch: submit in waves, poll, fetch results

A Python script that submits 30 Sume image jobs with mode async, keeps each SKU in metadata, polls the job status and reads artifacts from /v1/jobs/{id}/result.

5 min readSume
All posts

To run 30 image generations as one batch on Sume, submit each with mode: "async" so every call returns a 202 job envelope at once, poll GET /v1/jobs/{id}/status until terminal is true, then read the images from GET /v1/jobs/{id}/result. Keep your own id in metadata and send a stable Idempotency-Key per item so a retry never bills twice. The Image API docs and jobs and results define each step.

Do not submit all 30 at once on a small plan. The accepted capacity (processing plus queued) is 6 on Free and 24 on Pro, so the 25th Pro submit would fail with 429 queue_full. The script below submits in waves.

What does the script look like?

It uses requests, no top-level await, and prints each result envelope, whose artifacts[].url entries are the images. Each item uses Qwen Image at $0.025.

import os, time, requests

B = "https://api.sume.com/v1"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def submit(sku):
    r = requests.post(f"{B}/images", timeout=60,
        headers={**H, "Idempotency-Key": f"cover-{sku}"},
        json={"model": "qwen/qwen-image", "prompt": f"Cover for {sku}",
              "mode": "async", "metadata": {"sku": sku}})
    r.raise_for_status()
    return r.json()["data"]["job"]["id"]

def finish(job):
    while True:
        s = requests.get(f"{B}/jobs/{job}/status", headers=H, timeout=60).json()
        s = s.get("data", s)
        if s.get("terminal"):
            break
        time.sleep(s.get("next_poll_after_seconds") or 5)
    if s.get("sume_status") != "completed":
        return []
    return requests.get(f"{B}/jobs/{job}/result", headers=H, timeout=60).json()

skus = [f"sku-{i:03d}" for i in range(30)]
for i in range(0, len(skus), 18):
    wave = [submit(s) for s in skus[i:i + 18]]
    print([finish(j) for j in wave])

Why a wave of 18?

Sume's generation_limits snapshot carries a wave_size_hint: max(1, floor(queue_capacity_remaining x 0.75)). On an idle Pro workspace that is floor(24 x 0.75) = 18. It is a pacing hint, not a concurrency limit, so read the live value from the snapshot rather than hard-coding 18. See generation admission.

What should I store?

Store the job id next to your SKU. If the script dies, the jobs keep running and billing, and you can resume by polling the saved ids instead of resubmitting. A client-side timeout never cancels a job. The artifact URLs in the result are Sume-hosted, so use those rather than any provider link.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume