MAI-Image-2.6 allows 6 requests a minute at tier 1; Sume queues

Foundry rates MAI-Image-2.6 at 6 RPM on tier 1 and 429s past it. Sume accepts valid image jobs as queued until a plan slot opens. Compare the two behaviours.

5 min readSume
All posts

Microsoft Foundry limits MAI-Image-2.6 to 6 requests per minute on tier 1, rising by 6 per tier to 36 on tier 6; tier 0, the free tier, is 0 (Microsoft Learn, read 2026-10-01). Past the limit you get 429 Too Many Requests and must wait or request more quota. Sume treats a full workspace differently: valid image jobs are accepted as queued and start when a concurrency slot opens, and 429 queue_full appears only when the queue is also full.

That makes Sume's behaviour friendlier to a batch of 50 images, but not unlimited. Concurrency depends on the plan, not on top-ups, and the status endpoint is the source of truth.

How do the two limits differ?

One counts requests per minute; the other counts jobs in flight.

Rate behaviour, read 2026-10-01 from Microsoft Learn and Sume's generation admission docs
ItemMAI-Image-2.6 on FoundrySume Image API
UnitRequests per minute (RPM)Jobs processing at once, plus a queue
Small plansTier 1: 6 RPMFree: 1 processing, 5 queued
Larger plansTier 6: 36 RPMScale: 20 processing, 100 queued
When full429 Too Many RequestsAccepted as queued; 429 queue_full when the queue is full
More capacityQuota increase request formPlan change or admin override

How do I submit a batch on Sume without hitting 429?

Use mode: "async" so each call returns a job envelope at once, then poll the status URL. A queued job is not a failure, so just keep polling. The script submits three jobs and waits for each:

import os
import time
import requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def submit(prompt):
    r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=60,
        json={"model": "openai/gpt-image-2.5", "prompt": prompt, "mode": "async"})
    r.raise_for_status()
    return r.json()["data"]

def wait(job):
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s.get("terminal"):
            return s
        time.sleep(s.get("next_poll_after_seconds") or 3)

jobs = [submit(f"flat icon of a {x}") for x in ("kettle", "lamp", "chair")]
for job in jobs:
    print(job["job"]["id"], wait(job).get("sume_status"))

Limits

If you do hit 429 rate_limited on Sume, back off using retry-after when present; the errors page separates that from queue_full. Do not retry paid submits without an Idempotency-Key on routes that document one. MAI is not in Sume's image list as of this post, so this compares two designs, not two ways to call MAI. Microsoft's tiers depend on subscription and deployment, per the Learn page, so read your own quota before planning a batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume