Sume ratelimit-limit differs on reads and writes: read the header

ratelimit-limit shows the budget the current call spent from, so a GET shows reads and a POST shows writes. Read the header; do not hard-code the plan table.

5 min readSume
All posts

Short answer

Every Sume response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, and the docs say the headers describe whichever budget the current request spent from. A GET reports the read budget and a POST reports the write budget, so the same key shows two different numbers. Read the header on each response instead of hard-coding the plan table, because the read multiple is a deployment setting.

The shipped defaults

Each API key has a per-minute request budget set by the subscription plan of its workspace. Reads and writes have separate budgets, so a tight status-poll loop cannot 429 your own submits. A read is any GET or HEAD, plus POST /v1/generation/admission-preview and the MCP endpoint. Everything else is a write. The plan number is the write number, and reads get forty times that.

Per-key requests per minute by plan, shipped defaults (read 2026-10-03)
PlanWrites per minuteReads per minute
Free1204800
Pro30012000
Startup60024000
Scale120048000
EnterpriseContact salesContact sales

Why the header beats the table

The docs state that the read multiple is a deployment setting, SUME_COM_API_RATE_LIMIT_READ_MULTIPLIER, so a self-hosted or preview deployment can differ from the numbers above. ratelimit-limit on the response is always the authority for the deployment you are talking to. Enterprise has no self-serve number; until a contracted one is provisioned, an Enterprise key resolves to the Scale row.

Unauthenticated callers are limited per client IP at the Free rate, with the read bucket held at four times the write rate rather than forty. If you call a public route without a key, expect a much smaller read allowance than an authenticated key gets.

Using the headers

Use ratelimit-remaining to throttle before you hit the wall, and retry-after on a 429 to wait the right time. The 429 body names the exhausted budget in error.details.scope, either read or write. Do not confuse this with queue_full, which is also a 429 but means generation concurrency plus queue capacity is full; it is governed by the plan's concurrency limit and reported on the generation_limits object, and raising your request rate does not change it.

  • Log ratelimit-limit alongside the method, to see which bucket a call used.
  • Back off on retry-after rather than a fixed delay.
  • Alert on a falling ratelimit-remaining, not only on 429s.
  • Treat rate_limited and queue_full as separate causes with separate fixes.

A quick check

Call a read route and a write-shaped route with the same key and compare the two ratelimit-limit values. On a default deployment their ratio should match the multiple in the table. If it does not, you are talking to a deployment with a different setting, and your pacing code should follow what the headers say. Keep this check in a monitoring script, not in the request path of your product.

Planning for growth

If you outgrow a plan's write number, the lever is the plan, not a trick with the client. If your reads are the pressure, check the loop first: a poller that sleeps for next_poll_after_seconds on each reply spends far fewer reads than one that spins. Concurrency is separate again. How many generations run at once depends on the plan's concurrency limit, and exceeding it with a full queue gives queue_full rather than rate_limited.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume