Sume ratelimit-limit: read the budget from the header, not a table

Plan numbers are 120 to 1200 writes a minute, but dev and self-hosted deployments can differ. Calibrate a Python client from ratelimit-limit at startup.

4 min readSume
All posts

It is tempting to put the plan table in your code: Free 120, Pro 300, Startup 600, Scale 1200 writes a minute. The Sume docs warn against relying on it. The read multiple is a deployment setting, named SUME_COM_API_RATE_LIMIT_READ_MULTIPLIER, so a self-hosted or preview deployment can differ. The page says that ratelimit-limit on the response is always the authority for the deployment that you call. The table shows the shipped default.

What changes between deployments

  • The read multiple, which is 40 times the write number by default. Unauthenticated requests by client IP use 4 times.
  • The tier of a particular key, which follows the plan of its workspace.
  • An Enterprise key, which uses the Scale row until Sume provisions a contracted number.

What the header gives you

A read request returns the read limit, so one call tells you the reads. A write returns the write limit. If you want the write number without spending a write, divide the read limit by the multiple, but only when you know the multiple. When you are unsure, send one real write and read its header.

Default limits from the docs, and what to read instead (read 2026-10-04)
PlanDefault writes per minuteDefault reads per minute
Free1204800
Pro30012000
Startup60024000
Scale120048000

Calibrate at startup

The client below reads the limit from the first response and sizes a polling interval from it. It makes one request, uses only the standard library, and stores nothing between runs.

import os, urllib.request

def read_limit(base=None):
    base = base or os.environ.get("SUME_BASE_URL", "https://api.sume.com")
    req = urllib.request.Request(base + "/v1/me", headers={"x-api-key": os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=15) as resp:
        return int(resp.headers["ratelimit-limit"])

def poll_interval(jobs, limit, window=60, share=0.5):
    """Seconds between polls so that `jobs` pollers use at most `share` of the read budget."""
    return max(2.0, jobs * window / (limit * share))

limit = read_limit()
print(limit, poll_interval(50, limit))  # 4800 reads: 2.0 s for 50 jobs

Use dev for tests

Development integrations use https://api.dev.sume.com. Do not assume it has the same numbers as production. Read the header there too, and keep the base URL in configuration.

Log it so an upgrade is visible

Emit the value as a gauge when the client starts, and again once an hour in a long-running worker. After a plan change, the new number appears without a code change. The plan tier is cached for a short time on the server side, so allow about a minute after an upgrade before you expect the header to move.

If the number is lower than you expect, check which workspace owns the key. The budget follows the plan of the workspace that owns the key, not the plan of the person who typed the command.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume