Which Sume API requests count against the read rate limit?

Every GET and HEAD is a read, and so are POST /v1/generation/admission-preview and the MCP endpoint. Reads have their own per-key bucket, 40x the write one.

5 min readSume
All posts

Any GET or HEAD counts as a read. Two POST routes count as reads as well, POST /v1/generation/admission-preview and the MCP endpoint, because neither of them creates paid work. Everything else that writes counts as a write. Rate limits are per API key per minute, and reads and writes have separate buckets, so a polling loop cannot use up the budget you need to submit.

This matters most for dashboards and agents. A status board that refreshes every few seconds for dozens of jobs uses reads only, and an agent that checks /v1/balance before each action does the same. Both stay in the large read bucket, and neither slows down your submits.

Per minute limits by plan

The read bucket is 40 times the write bucket. That ratio exists because status polling is the normal way to follow a job. The limits scale with plan, and the table shows the per minute values per key.

The numbers above are per key. Two keys in the same workspace have two buckets, so a noisy poller on one key does not hurt another. Do not use that to dodge limits on purpose. Use it to give a background reporter its own key, so its traffic cannot drain the budget of the key that serves your customers.

Per key rate limits per minute (read 2026-10-05)
PlanWritesReads
Free1204800
Pro30012000
Startup60024000
Scale120048000

Public routes need no key

Anonymous calls to the public routes, such as the catalog, health check and OpenAPI document, have a read bucket that is 4 times the anonymous write bucket. Those routes need no key at all, so a monitor can call GET /v1/health without a credential.

Read the headers

Watch the response headers instead of guessing. Each response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, and a 429 adds retry-after. The error body names error.details.scope, which is read or write, so you know which bucket you emptied. A 429 with code queue_full is a different thing. It means the generation queue is full, not that you polled too fast.

Log the scope of every 429. A count of write scope 429 responses tells you to slow your submits or to buy a bigger plan, while a count of read scope 429 responses tells you to poll less often or to move to webhooks.

A header aware pause

The helper below turns a set of response headers into a decision. It runs as is, and you can unit test it with canned headers.

A fixed pause is also fine when the headers are missing. Wait at least the floor you chose, and double it on repeated 429 responses, up to a ceiling. The SDK client does the same thing for you. It retries 408, 429 and 5xx, honours retry-after up to 60 seconds, and adds some jitter, and it retries a POST only when an Idempotency-Key is present.

def pause_seconds(headers, status, floor=1.0):
    h = {k.lower(): v for k, v in headers.items()}
    if status == 429 and "retry-after" in h:
        return max(floor, float(h["retry-after"]))
    remaining = int(h.get("ratelimit-remaining", "1"))
    reset = float(h.get("ratelimit-reset", "0"))
    if remaining <= 0:
        return max(floor, reset)
    return 0.0

print(pause_seconds({"Retry-After": "7"}, 429))
print(pause_seconds({"ratelimit-remaining": "0", "ratelimit-reset": "12"}, 200))
print(pause_seconds({"ratelimit-remaining": "55"}, 200))

Polling faster does not help

Poll backpressure is not generation concurrency. The docs separate four controls, which are generation concurrency, queue capacity, submit rate limits and balance. Reading status faster does not move a job from queued to processing. The job moves when a workspace slot opens. Use the next_poll_after_seconds value from the job responses, and the poll pace follows what the server suggests.

Spend reads on purpose

For many jobs at once, prefer a webhook with a slow poll as the backup. A webhook sends one request per terminal event and uses none of your read budget. The timeline: true option of the SDK doubles the request rate of a run watcher, because each poll also reads the phase timeline, so turn it on only when you show phases to a person.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume