Upgraded your Sume plan but ratelimit-limit is still the old number?
A plan change can take up to 60 seconds to reach the per-key rate limit, because the tier is cached. Why ratelimit-limit lags, and what changes at once.

A plan change reaches the per-key request limit in about a minute, not instantly: the plan behind ratelimit-limit is cached with a 60-second default time to live, so the header can show the old number for roughly that long after you upgrade. The code keeps two such caches (one by workspace, one by key), so allow a little more than one minute. If it still shows the old value after a few minutes, check which key you are calling with before you suspect the cache.
This is easy to misread because two different limits both respond to a plan. One is the request budget per minute, which is what ratelimit-limit reports. The other is generation capacity, the number of jobs running at once. Only the first sits behind the 60-second cache. The facts here are from apps/api/src/rate-limit-tiers.ts, the Authentication rate-limit section and the Generation admission page, read 2026-10-10.
In short: wait a couple of minutes, call any /v1 route with the same key, and read the header again. The value is refreshed by authenticated requests that follow the cache expiring, so an idle key shows the old number until you use it.
Which number the header reports
Each API key gets a request budget per minute across /v1, set by the plan of the workspace that owns the key. Writes and reads have separate buckets. The table shows the shipped plan numbers; the Authentication page says the read multiple is deployment configuration and that ratelimit-limit on the response is the authority for the deployment you call.
Note that the table gives defaults. A workspace that Sume operations has given a concurrency override can have a request budget above its plan, equal to the seats the override grants per minute, per the rate-limit tier code. That is set by Sume, not bought, so it is not something an upgrade changes.
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4,800 |
| Pro | 300 | 12,000 |
| Startup | 600 | 24,000 |
| Scale | 1,200 | 48,000 |
| Enterprise | Contact sales | Contact sales |
Why the number lags
The limiter runs before authentication, on purpose, so a flood of invalid keys is stopped before it reaches the auth path. That means the limiter does not yet know whose plan to apply. The API resolves the plan after a request authenticates and writes it into a small cache that the limiter reads on the next request.
The cache time to live is DEFAULT_PLAN_TIER_CACHE_TTL_MS, 60,000 ms. The code comment calls it an upgrade-latency knob, not a correctness one: a plan change takes effect within one TTL, and until then the caller keeps the limit they had. Two safe fallbacks follow from the same design. A key the limiter has never seen is judged at the Free floor for one window. A plan the API cannot read also falls to the Free floor, never to a higher tier.
What a plan change does not delay
The request budget is not generation capacity. The code states that real throughput is owned by the workspace concurrency limit and the admission path, and that the two are deliberately not tied one to one, so a customer cannot buy past provider capacity by sending more requests. To see the capacity your workspace has right now, read generation_limits from the admission fields instead of inferring it from ratelimit-limit.
The reverse confusion also happens. A 429 rate_limited means the request budget is spent, and a 429 queue_full means the accepted-job capacity is spent. Retrying harder fixes neither faster. Wait for the retry-after seconds on a rate limit, and wait for jobs to finish or cancel queued ones on queue_full.
Watch the header change
The script below reads GET /v1/me every five seconds for three minutes and prints a line only when ratelimit-limit changes. Start it just before you upgrade. Reads use the generous read bucket, so it will not touch your submit budget. If the value never changes, confirm that the key belongs to the workspace you upgraded, since keys are workspace-scoped.
import os, time, urllib.request
def limit_header(key):
req = urllib.request.Request(
"https://api.sume.com/v1/me",
headers={"x-api-key": key},
)
with urllib.request.urlopen(req, timeout=10) as resp:
return resp.headers.get("ratelimit-limit")
def main():
key = os.environ.get("SUME_API_KEY", "")
if not key:
raise SystemExit("Set SUME_API_KEY first")
start = time.monotonic()
seen = None
while time.monotonic() - start < 180:
value = limit_header(key)
if value != seen:
print(f"{time.monotonic() - start:5.1f}s ratelimit-limit {value}")
seen = value
time.sleep(5)
main()What to do while you wait
Nothing needs to be rebuilt. Keep honoring retry-after on any 429, keep an Idempotency-Key on paid submits so a retry returns the original job, and pace a batch from generation_limits rather than from the header. After the TTL, the header reports the new tier. The page Cold key burst covers the same Free-floor rule for a brand-new key.
Sources
Related posts
More in Developers
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
- Spend caps for unattended AI agents: how Sume bounds each run
An unattended agent has no one to approve spend, so Sume caps generation per run: required on Agent Completions, and up to $500 on Format runs.
Written by Sume