Three Sume throttle signals: rate_limited, queue_full, 503

OpenAI now splits 429 (traffic rising too fast) from 503 (overload). Sume has three: 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded.

4 min readSume
All posts

OpenAI's changelog for September 2, 2026 says API errors now distinguish traffic increasing "too quickly" (429) from temporary model overload (503). Sume has three throttle-like signals: 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded. Each needs a different response from your client.

The OpenAI change

From the OpenAI API changelog, read on 2026-10-03: a 429 now means traffic is increasing too quickly, and a 503 means temporary model overload. Sume's codes are its own and are described below.

What each Sume signal means

The error body carries a code and a request id. Concurrency being full is not an error by itself: Sume accepts valid jobs as queued while queue capacity remains.

Sume throttle signals (per docs, read 2026-10-03)
Status and codeMeaningClient behavior
429 rate_limitedToo many requests in the current windowBack off, use retry-after when present
429 queue_fullWorkspace concurrency plus queue capacity is fullWait for a job to finish or cancel one
503 provider_capacity_exceededSume's provider dispatch queue is fullRetry later with the same idempotency key

Why queue_full is not a rate limit

queue_full means Sume cannot accept another paid job for the workspace until a queued or processing job finishes or is canceled. Sending requests more slowly does not fix it; finishing work does. rate_limited is about request volume, including polling.

One retry function

Never retry an unsafe submit without an Idempotency-Key, and reuse the same key across attempts so a retry returns the original job.

import time
import requests

def submit(url, headers, body, key, tries=5):
    headers = {**headers, "Idempotency-Key": key}
    for attempt in range(tries):
        r = requests.post(url, headers=headers, json=body, timeout=30)
        if r.status_code < 400:
            return r.json()
        code = r.json().get("error", {}).get("code")
        retry_after = float(r.headers.get("retry-after", 2 ** attempt))
        if code == "queue_full":
            retry_after = max(retry_after, 30)
        elif code not in ("rate_limited", "provider_capacity_exceeded"):
            r.raise_for_status()
        time.sleep(retry_after)
    raise RuntimeError("still throttled")

Log the request id

Record error.request_id with every throttled response. It is safe to share with Sume support. Do not log API keys or signed URLs.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume