Sume ratelimit-limit: read the budget from the header, not a table
Plan numbers are 120 to 1200 writes a minute, but dev and self-hosted deployments can differ. Calibrate a Python client from ratelimit-limit at startup.

It is tempting to put the plan table in your code: Free 120, Pro 300, Startup 600, Scale 1200 writes a minute. The Sume docs warn against relying on it. The read multiple is a deployment setting, named SUME_COM_API_RATE_LIMIT_READ_MULTIPLIER, so a self-hosted or preview deployment can differ. The page says that ratelimit-limit on the response is always the authority for the deployment that you call. The table shows the shipped default.
What changes between deployments
- The read multiple, which is 40 times the write number by default. Unauthenticated requests by client IP use 4 times.
- The tier of a particular key, which follows the plan of its workspace.
- An Enterprise key, which uses the Scale row until Sume provisions a contracted number.
What the header gives you
A read request returns the read limit, so one call tells you the reads. A write returns the write limit. If you want the write number without spending a write, divide the read limit by the multiple, but only when you know the multiple. When you are unsure, send one real write and read its header.
| Plan | Default writes per minute | Default reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
Calibrate at startup
The client below reads the limit from the first response and sizes a polling interval from it. It makes one request, uses only the standard library, and stores nothing between runs.
import os, urllib.request
def read_limit(base=None):
base = base or os.environ.get("SUME_BASE_URL", "https://api.sume.com")
req = urllib.request.Request(base + "/v1/me", headers={"x-api-key": os.environ["SUME_API_KEY"]})
with urllib.request.urlopen(req, timeout=15) as resp:
return int(resp.headers["ratelimit-limit"])
def poll_interval(jobs, limit, window=60, share=0.5):
"""Seconds between polls so that `jobs` pollers use at most `share` of the read budget."""
return max(2.0, jobs * window / (limit * share))
limit = read_limit()
print(limit, poll_interval(50, limit)) # 4800 reads: 2.0 s for 50 jobsUse dev for tests
Development integrations use https://api.dev.sume.com. Do not assume it has the same numbers as production. Read the header there too, and keep the base URL in configuration.
Log it so an upgrade is visible
Emit the value as a gauge when the client starts, and again once an hour in a long-running worker. After a plan change, the new number appears without a code change. The plan tier is cached for a short time on the server side, so allow about a minute after an upgrade before you expect the header to move.
If the number is lower than you expect, check which workspace owns the key. The budget follows the plan of the workspace that owns the key, not the plan of the person who typed the command.
Sources
Related posts
More in Developers
- Sume ratelimit-reset: sleep until the 60-second window ends (Python)
Every Sume /v1 response carries ratelimit-remaining and ratelimit-reset. Stop at zero and sleep the reset seconds instead of eating a 429. A Python wrapper.
- Does a second Sume API key raise your rate limit? No, here is why
Each Sume API key has its own bucket, but the workspace owner has an account bucket too. A second key spreads load without adding requests per minute.
- Authenticate the Sume CLI on a CI runner without a browser login
On CI, skip sume login: install the CLI, run sume auth setup with an API key from a secret, and confirm with sume auth status before any job step.
- sume/auto for a former Sora feature: when to pin a model
Sume's sume/auto picks a family and never says which. Good for general clips, wrong when a brand needs one look. How to choose between auto and a pinned id.
Written by Sume