429 rate_limited: how to tell a read budget from a write budget

A Sume 429 rate_limited names its budget in error.details.scope, read or write. Reads get forty times the write number. Headers to pace on and what to back off.

4 min readSume
All posts

A Sume 429 rate_limited names the budget it came from in error.details.scope, which is read or write. Reads and writes have separate per-minute budgets, and reads get forty times the write number, so a tight status-poll loop cannot 429 your own submits. Read the scope first, then back off on retry-after.

This follows Authentication and Format API errors, read 2026-09-29. The per-minute numbers depend on your plan, so read ratelimit-limit on a real response rather than trusting a table.

What counts as a read and what counts as a write?

A read is any GET or HEAD: polling status_url, events_url or result_url, and listing Formats or runs. The MCP endpoint itself also counts as a read, even though it takes POSTs. Everything else is a write: creating runs, cancelling, and so on.

An MCP tool call spends the write budget for the run it creates, once, not for the JSON-RPC request that carried it. A jobs_status poll over MCP spends no write budget at all.

Which headers should I pace on?

Every response carries the current state, and the headers describe whichever budget the current request spent from.

From Authentication, read 2026-09-29.
HeaderMeaning
ratelimit-limitRequests allowed in the current window
ratelimit-remainingRequests left in the current window
ratelimit-resetSeconds until the window resets
retry-afterSeconds to wait, sent on 429

How do I read the scope from a 429?

The error envelope is the same one every Sume error uses, with the scope under details. A write scope on a Format create means the write budget for that key is spent. Slow down submits; polling is on a separate budget and is not the cause.

code=$(curl -sS -D headers.txt -o resp.json -w "%{http_code}" \
  https://api.sume.com/v1/jobs/job_123/status \
  -H "Authorization: Bearer $SUME_API_KEY")
if [ "$code" = "429" ]; then
  grep -i '^retry-after:' headers.txt
  python3 -c "import json; print(json.load(open('resp.json'))['error']['details']['scope'])"
fi

Is this the same as queue_full?

No. 429 queue_full is a separate code: the workspace's generation concurrency plus queue capacity is full. Request rate is not generation capacity, and raising your request rate does not raise the plan's concurrency limit, which is reported on the generation_limits object. The fix for queue_full is to wait for jobs to finish or cancel queued jobs, not to slow your polling.

What are the numbers?

The Authentication docs list writes per minute and reads per minute by plan. On the read date the Free plan is 120 writes and 4800 reads a minute, and the Scale plan is 1200 and 48000; Enterprise is contracted. A key gets the budget of the workspace it belongs to, and a self-hosted or preview deployment can differ, which is why ratelimit-limit is the authority.

For a job submit the generation-admission page describes the same 429 rate_limited as an abuse-protection limit on request volume, with the same client behavior: back off using retry-after when present.

What about unauthenticated requests?

They are limited per client IP at the Free rate, and their read bucket is held at four times the write rate instead of forty. The agent-sized read budget is for callers who own the jobs they are polling.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume