A 429 from the Sume API: read budget, write budget or queue_full?
Three different 429s share a status. Read error.code and error.details.scope to tell read, write and queue_full apart, then retry each the right way.

Check error.code first, then error.details.scope. rate_limited with scope read means your polling used the read budget. rate_limited with scope write means your submits, cancels or uploads used the write budget. queue_full is not a request-rate limit at all: the workspace has no accepted generation capacity left. Each needs a different response, and treating them the same either wastes time or drops work.
Three 429s, three meanings
All three arrive with HTTP 429, so branch on the body. The headers ratelimit-limit, ratelimit-remaining and ratelimit-reset describe the budget the current request spent from, and retry-after appears on a 429.
| Signal | Meaning | What to do |
|---|---|---|
rate_limited, details.scope: read | Polls, status reads and list calls went above the read budget | Wait retry-after, read again. The job keeps running; do not resubmit. |
rate_limited, details.scope: write | Creates, cancels or uploads went above the write budget | Wait retry-after, then resend with the same Idempotency-Key. |
queue_full | Running plus queued paid jobs reached the workspace's accepted capacity | Poll current jobs, cancel queued ones you no longer need, retry with the same key when capacity opens. |
How big the budgets are
Each API key gets a per-minute budget across /v1, set by the plan of the workspace that owns it. Reads get forty times the write number in a separate bucket. In requests per second, a minute of budget divided by 60 gives the steady rate: Free 120 writes per minute is 2 per second, Pro 300 is 5, Startup 600 is 10 and Scale 1,200 is 20.
Compare those to the accepted-job ceilings: 6 on Free, 24 on Pro, 48 on Startup and 120 on Scale. Filling the whole accepted capacity in a minute costs 6, 24, 48 and 120 submits, which is 5%, 8%, 8% and 10% of the write budget. In practice a burst of submits meets queue_full long before it meets the write limit, so a 429 on a submit is more often queue_full than rate_limited. Read the code and you will know.
A classifier you can reuse
This function maps a response to a next step. It treats queue_full first because it needs a different action, then splits the rate limit on scope. Replace the return strings with your own scheduling calls.
export function nextStep(res, body) {
const e = body.error ?? {};
if (res.status !== 429) return "not-a-429";
if (e.code === "queue_full") return "wait-for-a-job-to-finish";
const wait = Number(res.headers.get("retry-after") ?? 5);
return e.details?.scope === "read"
? `poll-again-in-${wait}s`
: `resend-same-key-in-${wait}s`;
}Logging the right fields
When a 429 reaches your logs, record four things: the HTTP status, error.code, error.details.scope if present, and request_id. The request id is also sent as a response header and is safe to share with Sume support, unlike API keys, signed URLs or raw media URLs, which you should redact. A log line with those four fields lets you chart how many of your 429s are read, write or capacity, and that chart tells you what to change.
If most of them are read 429s, move long jobs to webhooks and slow the poll interval. If they are write 429s, spread your submits or move to a plan with a larger write budget. If they are queue_full, size your waves from the generation_limits object on the submit response, which reports the effective concurrency, the remaining queue capacity, and a wave_size_hint for how much to send next.
Mistakes that cost money
Do not resubmit a paid job because a status poll got a 429. The job exists and keeps billing; only your read failed. Do not retry an unsafe submit without an Idempotency-Key, because a retry that lands after the first one succeeded starts a second job. And do not count requests yourself to stay under the limit: read ratelimit-remaining from the response, which describes the deployment you are actually calling.
Separate the two buckets in your code, too. A tight status-poll loop cannot cause a write 429, because reads and writes are counted apart. If your submits are being limited, the cause is the submits. Enterprise is not self-serve; until a contracted number is provisioned, an Enterprise key uses the Scale row of the budget table, so plan for 1,200 writes per minute there.
Sources
Related posts
More in Developers
- Use your ad variant id as the Idempotency-Key on Sume
Name each ad variant once and send that name as the Idempotency-Key. Retries do not double-bill, and the receipt echoes it back so results map to your ad names.
- API key scopes for Sume: which key can call which endpoint family?
Sume API keys carry fixed scopes: formats:write, actions:read, agent_completions:write, account:read. Which scope each route needs, and why old keys get a 403.
- Arabic speech to text API: Sume STT with language_code ar
Transcribe Arabic audio with Sume STT: send language_code ar, check the reported language, and review the text. $0.01 per audio minute.
- Avatar video webhook mode: what arrives and what to poll anyway
Use mode webhook for a Sume avatar video and Sume posts one terminal event: completed, failed or canceled. Payload, signature headers and the polling backup.
Written by Sume