Rate limit headers: what limit, remaining and reset mean

Rate limit headers report your request budget: the window's limit, what's left, and when it resets. What each means and how a client should pace itself.

5 min readSume
All posts

Rate limit headers are HTTP response headers that report your request budget: how many requests the current window allows, how many you have left, and how long until the window resets. When you run out, a 429 Too Many Requests response may carry a Retry-After header that says how long to wait. A client should read them on every response and slow down as the remaining count nears zero, instead of waiting for the 429.

The conventions come from the IETF's RateLimit header fields draft and RFC 6585; the worked example is Sume's API, from its Authentication docs. All were read on 2026-09-28.

What do ratelimit-limit, ratelimit-remaining and ratelimit-reset mean?

The IETF draft describes the common choice as three headers: the maximum number of requests allowed in the time window, the number of remaining requests in the current window, and the time remaining in the window, in seconds or as a timestamp. Here is one API's version, with the 429 header added:

From Authentication, read 2026-09-28. The first three are on every response.
HeaderMeaning on the Sume API
ratelimit-limitRequests allowed in the current window
ratelimit-remainingRequests left in the current window
ratelimit-resetSeconds until the window resets
retry-afterSeconds to wait, sent on 429

Is there a standard for rate limit headers?

Not a finished one. The draft's own survey lists X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset as commonly used names, plus variants with the window in the name such as x-ratelimit-remaining-minute. It calls the lack of standard headers a major interoperability issue, because each implementation gives the same names different meanings. So read each API's docs for the reset value before you use it: seconds and a timestamp look alike in a header.

The IETF HTTPAPI working group's answer is two new fields, RateLimit-Policy and RateLimit. Version 11, dated 23 May 2026, is still an Internet-Draft, not an RFC. In its example RateLimit: "default";r=50;t=30, r is the available quota and t the effective window in seconds.

How should a client use ratelimit-remaining?

  • Read it; don't count. Sume's docs say to read ratelimit-remaining rather than counting requests yourself, and to back off on retry-after.
  • Treat it as a hint, not a promise. The IETF draft says clients must not assume a positive available quota guarantees the next request will be served.
  • Know which budget it describes. On Sume, reads and writes have separate budgets. The headers describe whichever budget the current request spent from, and a 429 names it in error.details.scope (read or write).
  • Wait the stated time after a 429. RFC 6585 says a 429 may include Retry-After; Retry-After after a 429 covers the case where it is missing.
const res = await fetch("https://api.sume.com/v1/jobs/job_123/status", {
  headers: { Authorization: `Bearer ${process.env.SUME_API_KEY}` },
});
const remaining = Number(res.headers.get("ratelimit-remaining"));
const reset = Number(res.headers.get("ratelimit-reset"));
if (res.status === 429) {
  const wait = Number(res.headers.get("retry-after") ?? reset);
  await new Promise((r) => setTimeout(r, wait * 1000));
} else if (remaining === 0) {
  await new Promise((r) => setTimeout(r, reset * 1000));
}

Does a rate limit cap how many jobs run at once?

No, not on an API that separates the two. Sume's docs say request rate is not the same thing as generation capacity: how many generations run at once is governed by the plan's concurrency limit, reported on the generation_limits object, and raising the request rate does not raise it.

What limits does Sume put in these headers?

Each API key gets a per-minute budget across all of /v1, set by the workspace's plan, and reads get forty times the write number in their own bucket, so a tight status-poll loop can't 429 your own submits. The per-plan numbers are in Sume API errors and rate limits. The docs call the response's ratelimit-limit the authority for the deployment you are talking to, so size your client from the header, not from a table.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume