Three Sume throttle signals: rate_limited, queue_full, 503
OpenAI now splits 429 (traffic rising too fast) from 503 (overload). Sume has three: 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded.

OpenAI's changelog for September 2, 2026 says API errors now distinguish traffic increasing "too quickly" (429) from temporary model overload (503). Sume has three throttle-like signals: 429 rate_limited, 429 queue_full and 503 provider_capacity_exceeded. Each needs a different response from your client.
The OpenAI change
From the OpenAI API changelog, read on 2026-10-03: a 429 now means traffic is increasing too quickly, and a 503 means temporary model overload. Sume's codes are its own and are described below.
What each Sume signal means
The error body carries a code and a request id. Concurrency being full is not an error by itself: Sume accepts valid jobs as queued while queue capacity remains.
| Status and code | Meaning | Client behavior |
|---|---|---|
| 429 rate_limited | Too many requests in the current window | Back off, use retry-after when present |
| 429 queue_full | Workspace concurrency plus queue capacity is full | Wait for a job to finish or cancel one |
| 503 provider_capacity_exceeded | Sume's provider dispatch queue is full | Retry later with the same idempotency key |
Why queue_full is not a rate limit
queue_full means Sume cannot accept another paid job for the workspace until a queued or processing job finishes or is canceled. Sending requests more slowly does not fix it; finishing work does. rate_limited is about request volume, including polling.
One retry function
Never retry an unsafe submit without an Idempotency-Key, and reuse the same key across attempts so a retry returns the original job.
import time
import requests
def submit(url, headers, body, key, tries=5):
headers = {**headers, "Idempotency-Key": key}
for attempt in range(tries):
r = requests.post(url, headers=headers, json=body, timeout=30)
if r.status_code < 400:
return r.json()
code = r.json().get("error", {}).get("code")
retry_after = float(r.headers.get("retry-after", 2 ** attempt))
if code == "queue_full":
retry_after = max(retry_after, 30)
elif code not in ("rate_limited", "provider_capacity_exceeded"):
r.raise_for_status()
time.sleep(retry_after)
raise RuntimeError("still throttled")Log the request id
Record error.request_id with every throttled response. It is safe to share with Sume support. Do not log API keys or signed URLs.
Sources
Related posts
More in Developers
- Track AI video spend per job: Synthesia Billing API vs Sume job cost
Synthesia added a Billing API and auto top-up. On Sume, every finished job reports usage.cost, so you can keep a per-job ledger without a billing endpoint.
- Transcribe two minutes of a long video: audio detach range, then STT
Streaming transcribers charge by the hour; you may only need one segment. Detach a range as 16 kHz mono wav, then run one STT job. Caps and codes included.
- Trigger.dev Node 21 warning: which Node runs the Sume SDK
Trigger.dev v4.6.1 added Node.js 21 deprecation warnings. The Sume TypeScript SDK needs Node 18 or later, so tasks on Node 22 or newer are fine.
- Trigger.dev public tokens: keep the Sume key server-side
Trigger.dev v4.6.2 hardened authorization for public tokens. Whatever token your browser holds, a Sume API key must never be one of them. Here is the split.
Written by Sume