Cold Sume API key burst: first requests get the Free 120-write floor

The API reads your plan after auth, so a first-seen key is judged at the Free 120 write floor. Warm a Pro or Scale key with GET /v1/me before a burst.

4 min readSume
All posts

A paid key that has been idle can see a lower limit on its very first requests. The Sume API resolves the plan of a key after authentication. Before then, the per-key check has no tier for a credential it has not seen, so it uses the Free floor of 120 writes per minute. After the first authenticated request the plan is written through to a cache, and the key lands on its real tier.

What the code says

The design comment in the rate limiter states the rule: a first-seen credential is limited at the Free floor for one window and lands on its real tier from then on. The cache entry lasts 60 seconds, which equals the window, so a key that has been idle for a minute starts cold again. The floor is the value every caller already had, so a sequential client never notices.

When it matters

It matters when a Pro, Startup or Scale key opens with a burst of parallel POSTs, for example a nightly batch that starts 400 submits at once. Requests that are checked before the first one finishes are judged at 120, and the rest can return 429 rate_limited even though the plan allows 300 to 1200. The account bucket uses the real tier, so it is not the cause.

Warm the key first

A read is cheap, since reads have their own bucket at forty times the write number. One GET /v1/me before the burst is enough to write the tier through. Then check ratelimit-limit on a write before you fan out.

import os, urllib.request

def warm_and_report(base="https://api.sume.com"):
    req = urllib.request.Request(base + "/v1/me", headers={"x-api-key": os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=15) as resp:
        resp.read()
        # a read bucket: forty times the write number, 4800 on Free
        return int(resp.headers["ratelimit-limit"])

print("read limit:", warm_and_report())

Reading the result

The read limit shows your plan indirectly: 4800 is Free, 12000 is Pro, 24000 is Startup and 48000 is Scale. Divide by 40 for the write number. If the number does not match your plan, the key may belong to another workspace, so check the owner in the dashboard before you contact support.

Keep the batch honest

Warming does not replace a retry policy. Keep honoring retry-after on 429, and pace the wave to the write budget. Sustained load needs the key to stay warm, and any request in each 60 seconds does that.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume