Ideogram API 10 in-flight requests: how Sume queues image jobs instead
Ideogram's default limit is 10 in-flight requests. Sume accepts more jobs than it runs, by plan, and answers 429 queue_full only when the queue is full.

Ideogram's overview (read 2026-10-06) lists a default rate limit of 10 in-flight requests. That is a cap on how many calls you may have open at once.
Sume limits something different. Its admission docs say concurrency is a dispatch limit, not a submit limit: a workspace at its processing cap can still submit valid jobs, and they wait as queued. Only when the queue is also full does a paid submission fail with 429 queue_full.
The numbers on Sume
The docs tell you to prefer the effective generation_limits.concurrency_limit field over this static table, since admin overrides change it. The default queue capacity is max(3, concurrency_limit x 5).
| Plan | Processing concurrency | Queue capacity | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
What to change in a port
- Drop the client-side gate at 10. Gate at accepted job capacity instead, and submit with
mode: "async". - A
429 queue_fullis not a rate limit. Wait for a job to finish, then retry with the sameIdempotency-Key. - A normal
429 rate_limitedhasretry-after; use it when present. - Read results from
GET /v1/jobs/{id}/result, or use a webhook for the terminal event.
Submit a small wave
This sends six low-quality draft jobs without waiting, then prints each job id. On a Free workspace that is exactly the accepted capacity of 6.
import os, requests
key = os.environ.get("SUME_API_KEY")
if not key:
raise SystemExit("set SUME_API_KEY")
for i in range(6):
r = requests.post(
"https://api.sume.com/v1/images",
headers={"Authorization": f"Bearer {key}", "Idempotency-Key": f"wave-{i}"},
json={"model": "ideogram/ideogram-v4.5", "quality": "low", "mode": "async",
"prompt": f"Flat badge number {i + 1}, bold sans-serif"},
timeout=60,
)
if r.status_code == 429:
print(i, "429", r.json())
break
print(i, r.status_code, r.json()["data"]["job"]["id"])Reading your own limit
The table above is the default. Your workspace may differ, and the docs say the effective value is the generation_limits.concurrency_limit field. Read it at start-up, compute the accepted capacity as concurrency plus queue, and keep your in-flight submissions below that number.
If you need a bigger wave than your plan accepts, split it into batches and wait for the queue to drain between them. That is cheaper than retrying rejected submissions, because a queue_full response means nothing was accepted.
Sources
Related posts
More in Comparisons
- Ideogram API lists GPT Image and Nano Banana: what Sume's catalog adds
Ideogram's API overview now lists GPT Image and Nano Banana beside Ideogram 4.5. Here is how that compares with the models Sume's /v1/images catalog lists.
- Interactive AI video for enterprise: live avatar or quiz video?
Interactive can mean a live conversation or a quiz inside a video. A table of both from HeyGen's September guide, and which part a Sume avatar clip covers.
- Kling 4.0 or Seedance 2.5 for a 30-second AI video today?
Both advertise 30 seconds. Only one has a callable id on Sume today. A spec-by-spec read of the vendor pages and what you can actually request.
- Lemonfox TTS at $2.50 per million characters vs Sume: the honest gap
Lemonfox's page works out to about $2.50 per million characters. Sume lists $47.50. Cost for 900, 22,500 and 1M characters, and what the gap buys you.
Written by Sume