Upgraded your Sume plan but ratelimit-limit is still the old number?

A plan change can take up to 60 seconds to reach the per-key rate limit, because the tier is cached. Why ratelimit-limit lags, and what changes at once.

5 min readSume
All posts

A plan change reaches the per-key request limit in about a minute, not instantly: the plan behind ratelimit-limit is cached with a 60-second default time to live, so the header can show the old number for roughly that long after you upgrade. The code keeps two such caches (one by workspace, one by key), so allow a little more than one minute. If it still shows the old value after a few minutes, check which key you are calling with before you suspect the cache.

This is easy to misread because two different limits both respond to a plan. One is the request budget per minute, which is what ratelimit-limit reports. The other is generation capacity, the number of jobs running at once. Only the first sits behind the 60-second cache. The facts here are from apps/api/src/rate-limit-tiers.ts, the Authentication rate-limit section and the Generation admission page, read 2026-10-10.

In short: wait a couple of minutes, call any /v1 route with the same key, and read the header again. The value is refreshed by authenticated requests that follow the cache expiring, so an idle key shows the old number until you use it.

Which number the header reports

Each API key gets a request budget per minute across /v1, set by the plan of the workspace that owns the key. Writes and reads have separate buckets. The table shows the shipped plan numbers; the Authentication page says the read multiple is deployment configuration and that ratelimit-limit on the response is the authority for the deployment you call.

Note that the table gives defaults. A workspace that Sume operations has given a concurrency override can have a request budget above its plan, equal to the seats the override grants per minute, per the rate-limit tier code. That is set by Sume, not bought, so it is not something an upgrade changes.

Per-key requests per minute by plan, from the Authentication docs rate-limit table (read 2026-10-10)
PlanWrites per minuteReads per minute
Free1204,800
Pro30012,000
Startup60024,000
Scale1,20048,000
EnterpriseContact salesContact sales

Why the number lags

The limiter runs before authentication, on purpose, so a flood of invalid keys is stopped before it reaches the auth path. That means the limiter does not yet know whose plan to apply. The API resolves the plan after a request authenticates and writes it into a small cache that the limiter reads on the next request.

The cache time to live is DEFAULT_PLAN_TIER_CACHE_TTL_MS, 60,000 ms. The code comment calls it an upgrade-latency knob, not a correctness one: a plan change takes effect within one TTL, and until then the caller keeps the limit they had. Two safe fallbacks follow from the same design. A key the limiter has never seen is judged at the Free floor for one window. A plan the API cannot read also falls to the Free floor, never to a higher tier.

What a plan change does not delay

The request budget is not generation capacity. The code states that real throughput is owned by the workspace concurrency limit and the admission path, and that the two are deliberately not tied one to one, so a customer cannot buy past provider capacity by sending more requests. To see the capacity your workspace has right now, read generation_limits from the admission fields instead of inferring it from ratelimit-limit.

The reverse confusion also happens. A 429 rate_limited means the request budget is spent, and a 429 queue_full means the accepted-job capacity is spent. Retrying harder fixes neither faster. Wait for the retry-after seconds on a rate limit, and wait for jobs to finish or cancel queued ones on queue_full.

Watch the header change

The script below reads GET /v1/me every five seconds for three minutes and prints a line only when ratelimit-limit changes. Start it just before you upgrade. Reads use the generous read bucket, so it will not touch your submit budget. If the value never changes, confirm that the key belongs to the workspace you upgraded, since keys are workspace-scoped.

import os, time, urllib.request

def limit_header(key):
    req = urllib.request.Request(
        "https://api.sume.com/v1/me",
        headers={"x-api-key": key},
    )
    with urllib.request.urlopen(req, timeout=10) as resp:
        return resp.headers.get("ratelimit-limit")

def main():
    key = os.environ.get("SUME_API_KEY", "")
    if not key:
        raise SystemExit("Set SUME_API_KEY first")
    start = time.monotonic()
    seen = None
    while time.monotonic() - start < 180:
        value = limit_header(key)
        if value != seen:
            print(f"{time.monotonic() - start:5.1f}s  ratelimit-limit {value}")
            seen = value
        time.sleep(5)

main()

What to do while you wait

Nothing needs to be rebuilt. Keep honoring retry-after on any 429, keep an Idempotency-Key on paid submits so a retry returns the original job, and pace a batch from generation_limits rather than from the header. After the TTL, the header reports the new tier. The page Cold key burst covers the same Free-floor rule for a brand-new key.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume