MAI-Image-2.6 allows 6 requests a minute at tier 1; Sume queues
Foundry rates MAI-Image-2.6 at 6 RPM on tier 1 and 429s past it. Sume accepts valid image jobs as queued until a plan slot opens. Compare the two behaviours.

Microsoft Foundry limits MAI-Image-2.6 to 6 requests per minute on tier 1, rising by 6 per tier to 36 on tier 6; tier 0, the free tier, is 0 (Microsoft Learn, read 2026-10-01). Past the limit you get 429 Too Many Requests and must wait or request more quota. Sume treats a full workspace differently: valid image jobs are accepted as queued and start when a concurrency slot opens, and 429 queue_full appears only when the queue is also full.
That makes Sume's behaviour friendlier to a batch of 50 images, but not unlimited. Concurrency depends on the plan, not on top-ups, and the status endpoint is the source of truth.
How do the two limits differ?
One counts requests per minute; the other counts jobs in flight.
| Item | MAI-Image-2.6 on Foundry | Sume Image API |
|---|---|---|
| Unit | Requests per minute (RPM) | Jobs processing at once, plus a queue |
| Small plans | Tier 1: 6 RPM | Free: 1 processing, 5 queued |
| Larger plans | Tier 6: 36 RPM | Scale: 20 processing, 100 queued |
| When full | 429 Too Many Requests | Accepted as queued; 429 queue_full when the queue is full |
| More capacity | Quota increase request form | Plan change or admin override |
How do I submit a batch on Sume without hitting 429?
Use mode: "async" so each call returns a job envelope at once, then poll the status URL. A queued job is not a failure, so just keep polling. The script submits three jobs and waits for each:
import os
import time
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def submit(prompt):
r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=60,
json={"model": "openai/gpt-image-2.5", "prompt": prompt, "mode": "async"})
r.raise_for_status()
return r.json()["data"]
def wait(job):
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s.get("terminal"):
return s
time.sleep(s.get("next_poll_after_seconds") or 3)
jobs = [submit(f"flat icon of a {x}") for x in ("kettle", "lamp", "chair")]
for job in jobs:
print(job["job"]["id"], wait(job).get("sume_status"))Limits
If you do hit 429 rate_limited on Sume, back off using retry-after when present; the errors page separates that from queue_full. Do not retry paid submits without an Idempotency-Key on routes that document one. MAI is not in Sume's image list as of this post, so this compares two designs, not two ways to call MAI. Microsoft's tiers depend on subscription and deployment, per the Learn page, so read your own quota before planning a batch.
Sources
Related posts
More in Developers
- MAI-Image-2.6 returns base64 PNG; Sume returns a hosted image URL
Foundry returns MAI-Image-2.6 as b64_json PNG only. Sume returns a signed media URL in data[].url, in png, jpeg or webp per model. How to handle each in code.
- Make an AI avatar video from the terminal with the Sume CLI
sume avatars create and sume avatar-videos create submit Avatar 1.0 jobs from a shell. Flags, the --confirm-paid guard, and how to recover the job.
- Migrate Veo 3.1 API calls to Gemini Omni Flash 1.1: parameter map
Veo 3.1 previews shut down October 22, 2026. A parameter-by-parameter map from the Veo guide to Gemini Omni Flash 1.1 requests on Sume, with a working curl.
- Music API has no duration field: steer length in the prompt
Sume's music router rejects duration and duration_seconds. Ask for a 30-second or 2-minute track in the prompt, with section timestamps, and verify.
Written by Sume