Read Sume rate limit headers with curl: remaining, reset, retry-after

Run curl -D - on any /v1 route to read ratelimit-limit, ratelimit-remaining, ratelimit-reset, and retry-after on a 429. Anonymous reads are 4x Free writes.

5 min readSume
All posts

Add -D - and -o /dev/null to a curl call, and the response headers print without the body. Look for ratelimit-limit, ratelimit-remaining and ratelimit-reset on every response, and for retry-after on a 429. You can try it with no key, because GET /v1/health is a public route. Authenticated calls show the budget of your plan instead of the anonymous one.

The headers describe the budget that the current request spent from. That detail matters, because a read and a write draw from different buckets. A GET for the balance shows the read bucket, and a POST that creates work shows the write bucket, so check the headers on the type of call that you worry about.

Keep this habit for any new integration. Print the headers once from the machine that will run the job, because a shared egress address, a proxy or a different plan can change what you see.

What each header says

The four headers have fixed meanings in the authentication docs. ratelimit-limit is the number of requests permitted in the current window. ratelimit-remaining is what is left. ratelimit-reset is the number of seconds until the window resets. retry-after is the number of seconds to wait, and it is sent on a 429.

Do not count requests yourself. The docs give this advice directly, because your count and the server count drift apart as soon as there is a second process, a retry or a second key. The header is the source of truth for the deployment that you call, and the read multiple is a deployment setting, so a preview deployment can differ from the numbers in the docs.

What to expect by plan

Rate limits are per key per minute, with a write bucket and a separate read bucket. The read bucket is 40 times the write bucket for keyed accounts.

Treat the numbers as a floor for planning, not as a promise. The read multiple is a deployment setting, so a self-hosted or preview deployment can use other values. The ratelimit-limit header on the response is always the authority for the deployment you call, and a script that adapts to the header keeps working when a limit changes.

One more point for agents and automations. An MCP tool call spends the write budget once, for the run that it creates, and it does not spend the write budget for the JSON-RPC request that carried it, and a status poll over MCP spends no write budget at all.

Per key limits per minute by plan (read 2026-10-05)
PlanWrite bucketRead bucket
Free1204800
Pro30012000
Startup60024000
Scale120048000
Unauthenticated, per client IPFree write rate4 times the write rate

Two curl commands

Run the two commands below. The first reads the public health route with no key and prints only the headers that matter. The second does the same for the balance route with your key, and it needs the key in an environment variable.

The anonymous example needs no key, so you can run it anywhere. Unauthenticated calls are limited per client IP at the Free rate, and their read bucket is four times the write rate, not forty, so the figure you see is far lower than the figure for a keyed account.

A small script can sleep until the reset when ratelimit-remaining reaches zero. That is gentler than waiting for the 429, and it keeps your own logs free of errors that were avoidable. When you do get a 429, prefer retry-after to your own guess.

# no key needed: a public route
curl -s -D - -o /dev/null https://api.sume.com/v1/health | grep -i -E '^(ratelimit|retry-after|x-sume-request-id)'

# with a key: pick ONE credential header, never both
curl -s -D - -o /dev/null https://api.sume.com/v1/balance \
  -H "x-api-key: $SUME_API_KEY" | grep -i -E '^(ratelimit|retry-after|x-sume-request-id)'

When you get a 429

On a 429, read error.details.scope in the body. It is read or write, and it tells you which bucket you emptied. The code is rate_limited for request volume, and the docs say to use retry-after for the backoff when it is present. A 429 with the code queue_full is different. It means the generation queue is full, and polling more slowly will not fix it.

The x-sume-request-id header is worth capturing as well. It is exposed next to the rate limit headers, and it is the id to quote when you ask support about one call.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume