Cold Sume API key burst: first requests get the Free 120-write floor
The API reads your plan after auth, so a first-seen key is judged at the Free 120 write floor. Warm a Pro or Scale key with GET /v1/me before a burst.

A paid key that has been idle can see a lower limit on its very first requests. The Sume API resolves the plan of a key after authentication. Before then, the per-key check has no tier for a credential it has not seen, so it uses the Free floor of 120 writes per minute. After the first authenticated request the plan is written through to a cache, and the key lands on its real tier.
What the code says
The design comment in the rate limiter states the rule: a first-seen credential is limited at the Free floor for one window and lands on its real tier from then on. The cache entry lasts 60 seconds, which equals the window, so a key that has been idle for a minute starts cold again. The floor is the value every caller already had, so a sequential client never notices.
When it matters
It matters when a Pro, Startup or Scale key opens with a burst of parallel POSTs, for example a nightly batch that starts 400 submits at once. Requests that are checked before the first one finishes are judged at 120, and the rest can return 429 rate_limited even though the plan allows 300 to 1200. The account bucket uses the real tier, so it is not the cause.
Warm the key first
A read is cheap, since reads have their own bucket at forty times the write number. One GET /v1/me before the burst is enough to write the tier through. Then check ratelimit-limit on a write before you fan out.
import os, urllib.request
def warm_and_report(base="https://api.sume.com"):
req = urllib.request.Request(base + "/v1/me", headers={"x-api-key": os.environ["SUME_API_KEY"]})
with urllib.request.urlopen(req, timeout=15) as resp:
resp.read()
# a read bucket: forty times the write number, 4800 on Free
return int(resp.headers["ratelimit-limit"])
print("read limit:", warm_and_report())Reading the result
The read limit shows your plan indirectly: 4800 is Free, 12000 is Pro, 24000 is Startup and 48000 is Scale. Divide by 40 for the write number. If the number does not match your plan, the key may belong to another workspace, so check the owner in the dashboard before you contact support.
Keep the batch honest
Warming does not replace a retry policy. Keep honoring retry-after on 429, and pace the wave to the write budget. Sustained load needs the key to stay warm, and any request in each 60 seconds does that.
Sources
Related posts
More in Developers
- Sume ratelimit-limit: read the budget from the header, not a table
Plan numbers are 120 to 1200 writes a minute, but dev and self-hosted deployments can differ. Calibrate a Python client from ratelimit-limit at startup.
- Sume ratelimit-reset: sleep until the 60-second window ends (Python)
Every Sume /v1 response carries ratelimit-remaining and ratelimit-reset. Stop at zero and sleep the reset seconds instead of eating a 429. A Python wrapper.
- Does a second Sume API key raise your rate limit? No, here is why
Each Sume API key has its own bucket, but the workspace owner has an account bucket too. A second key spreads load without adding requests per minute.
- Sume write limits as submits per second: a Python pacer by plan
Free 120, Pro 300, Startup 600, Scale 1200 writes per minute is 2, 5, 10 and 20 per second. A fixed-window pacer in Python that never trips the 429.
Written by Sume