Rate limit headers: what limit, remaining and reset mean
Rate limit headers report your request budget: the window's limit, what's left, and when it resets. What each means and how a client should pace itself.

Rate limit headers are HTTP response headers that report your request budget: how many requests the current window allows, how many you have left, and how long until the window resets. When you run out, a 429 Too Many Requests response may carry a Retry-After header that says how long to wait. A client should read them on every response and slow down as the remaining count nears zero, instead of waiting for the 429.
The conventions come from the IETF's RateLimit header fields draft and RFC 6585; the worked example is Sume's API, from its Authentication docs. All were read on 2026-09-28.
What do ratelimit-limit, ratelimit-remaining and ratelimit-reset mean?
The IETF draft describes the common choice as three headers: the maximum number of requests allowed in the time window, the number of remaining requests in the current window, and the time remaining in the window, in seconds or as a timestamp. Here is one API's version, with the 429 header added:
| Header | Meaning on the Sume API |
|---|---|
ratelimit-limit | Requests allowed in the current window |
ratelimit-remaining | Requests left in the current window |
ratelimit-reset | Seconds until the window resets |
retry-after | Seconds to wait, sent on 429 |
Is there a standard for rate limit headers?
Not a finished one. The draft's own survey lists X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset as commonly used names, plus variants with the window in the name such as x-ratelimit-remaining-minute. It calls the lack of standard headers a major interoperability issue, because each implementation gives the same names different meanings. So read each API's docs for the reset value before you use it: seconds and a timestamp look alike in a header.
The IETF HTTPAPI working group's answer is two new fields, RateLimit-Policy and RateLimit. Version 11, dated 23 May 2026, is still an Internet-Draft, not an RFC. In its example RateLimit: "default";r=50;t=30, r is the available quota and t the effective window in seconds.
How should a client use ratelimit-remaining?
- Read it; don't count. Sume's docs say to read
ratelimit-remainingrather than counting requests yourself, and to back off onretry-after. - Treat it as a hint, not a promise. The IETF draft says clients must not assume a positive available quota guarantees the next request will be served.
- Know which budget it describes. On Sume, reads and writes have separate budgets. The headers describe whichever budget the current request spent from, and a
429names it inerror.details.scope(readorwrite). - Wait the stated time after a 429. RFC 6585 says a 429 may include
Retry-After; Retry-After after a 429 covers the case where it is missing.
const res = await fetch("https://api.sume.com/v1/jobs/job_123/status", {
headers: { Authorization: `Bearer ${process.env.SUME_API_KEY}` },
});
const remaining = Number(res.headers.get("ratelimit-remaining"));
const reset = Number(res.headers.get("ratelimit-reset"));
if (res.status === 429) {
const wait = Number(res.headers.get("retry-after") ?? reset);
await new Promise((r) => setTimeout(r, wait * 1000));
} else if (remaining === 0) {
await new Promise((r) => setTimeout(r, reset * 1000));
}Does a rate limit cap how many jobs run at once?
No, not on an API that separates the two. Sume's docs say request rate is not the same thing as generation capacity: how many generations run at once is governed by the plan's concurrency limit, reported on the generation_limits object, and raising the request rate does not raise it.
What limits does Sume put in these headers?
Each API key gets a per-minute budget across all of /v1, set by the workspace's plan, and reads get forty times the write number in their own bucket, so a tight status-poll loop can't 429 your own submits. The per-plan numbers are in Sume API errors and rate limits. The docs call the response's ratelimit-limit the authority for the deployment you are talking to, so size your client from the header, not from a table.
Sources
Related posts
More in Developers
- Retryable HTTP status codes: which errors to retry
Retry network errors, 408, 429, and 5xx with backoff; skip most other 4xx. Retry a POST only with an idempotency key, and read the API's retry flag.
- Speech to text API in Java: transcribe audio with HttpClient
Call a speech to text API from Java with the JDK HttpClient and Jackson: post the audio URL, poll the job, then read the transcript and word times.
- Stateless MCP server: sessions, Mcp-Session-Id, and handles
A stateless MCP server keeps no session between requests. What Mcp-Session-Id does, what 2026-07-28 removed, and where state goes instead.
- Synthesia API documentation: the quickstart, step by step
Synthesia's API docs make a video in four steps: a Legacy (v2) key, POST /v2/videos, poll until complete, then download from a time-limited link.
Written by Sume