Ideogram API 10 in-flight requests: how Sume queues image jobs instead

Ideogram's default limit is 10 in-flight requests. Sume accepts more jobs than it runs, by plan, and answers 429 queue_full only when the queue is full.

4 min readSume
All posts

Ideogram's overview (read 2026-10-06) lists a default rate limit of 10 in-flight requests. That is a cap on how many calls you may have open at once.

Sume limits something different. Its admission docs say concurrency is a dispatch limit, not a submit limit: a workspace at its processing cap can still submit valid jobs, and they wait as queued. Only when the queue is also full does a paid submission fail with 429 queue_full.

The numbers on Sume

The docs tell you to prefer the effective generation_limits.concurrency_limit field over this static table, since admin overrides change it. The default queue capacity is max(3, concurrency_limit x 5).

Default plan capacity from Sume's generation admission page, read 2026-10-06
PlanProcessing concurrencyQueue capacityAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120

What to change in a port

  • Drop the client-side gate at 10. Gate at accepted job capacity instead, and submit with mode: "async".
  • A 429 queue_full is not a rate limit. Wait for a job to finish, then retry with the same Idempotency-Key.
  • A normal 429 rate_limited has retry-after; use it when present.
  • Read results from GET /v1/jobs/{id}/result, or use a webhook for the terminal event.

Submit a small wave

This sends six low-quality draft jobs without waiting, then prints each job id. On a Free workspace that is exactly the accepted capacity of 6.

import os, requests

key = os.environ.get("SUME_API_KEY")
if not key:
    raise SystemExit("set SUME_API_KEY")
for i in range(6):
    r = requests.post(
        "https://api.sume.com/v1/images",
        headers={"Authorization": f"Bearer {key}", "Idempotency-Key": f"wave-{i}"},
        json={"model": "ideogram/ideogram-v4.5", "quality": "low", "mode": "async",
              "prompt": f"Flat badge number {i + 1}, bold sans-serif"},
        timeout=60,
    )
    if r.status_code == 429:
        print(i, "429", r.json())
        break
    print(i, r.status_code, r.json()["data"]["job"]["id"])

Reading your own limit

The table above is the default. Your workspace may differ, and the docs say the effective value is the generation_limits.concurrency_limit field. Read it at start-up, compute the accepted capacity as concurrency plus queue, and keep your in-flight submissions below that number.

If you need a bigger wave than your plan accepts, split it into batches and wait for the queue to drain between them. That is cheaper than retrying rejected submissions, because a queue_full response means nothing was accepted.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume