Does a second Sume API key raise your rate limit? No, here is why
Each Sume API key has its own bucket, but the workspace owner has an account bucket too. A second key spreads load without adding requests per minute.

Creating a second API key does not give you a second request budget. The Sume API checks two buckets on every /v1 request: one per credential, before auth, and one per account, after auth. The account bucket is keyed on the workspace and the owning user, so two keys from the same owner share it.
The two checks
The per-key check protects the API from invalid or abusive credentials, and it runs first. The per-account check runs once the key is resolved, using the plan tier of the workspace. Both use a fixed 60-second window, and both split reads from writes. The 429 you receive names the budget that refused it in error.details.scope (read or write).
| Bucket | Keyed on | Checked | Limit |
|---|---|---|---|
| Per key | The credential hash | Before auth | Plan tier, Free floor until the plan is resolved |
| Per account | Workspace and owner user | After auth | Plan tier of the workspace |
Numbers
The write budget per minute is Free 120, Pro 300, Startup 600 and Scale 1200. Reads get forty times that, so Free reads are 4800 per minute. A key that sends 100 writes a minute and a second key that sends 100 more from the same owner make 200 writes against one 120-write account bucket on the Free plan, and the account bucket returns 429 on the 121st.
What a second key is still good for
- Revoking one integration without breaking the others.
- Separating per-service logs, since each key shows up separately in your own telemetry.
- Keeping a failing key's pre-auth bucket from refusing healthy keys.
Check which bucket fired
The headers describe the budget that the current request spent from, so compare ratelimit-limit with your plan row. A small Python probe makes the shared budget visible.
import os, urllib.request
def budget(key):
req = urllib.request.Request("https://api.sume.com/v1/me", headers={"x-api-key": key})
with urllib.request.urlopen(req, timeout=15) as resp:
return {h: resp.headers.get(h) for h in ("ratelimit-limit", "ratelimit-remaining", "ratelimit-reset")}
for name in ("SUME_API_KEY_A", "SUME_API_KEY_B"):
print(name, budget(os.environ[name]))When you need more
Raise the plan, not the key count. The plan tier sets the write number, and the docs say Enterprise uses the Scale row until a contracted number is provisioned. Request rate is also separate from concurrency: the generation_limits object reports how many generations can run at once, and more requests per minute do not raise it.
Sources
Related posts
More in Developers
- Sume write limits as submits per second: a Python pacer by plan
Free 120, Pro 300, Startup 600, Scale 1200 writes per minute is 2, 5, 10 and 20 per second. A fixed-window pacer in Python that never trips the 429.
- Authenticate the Sume CLI on a CI runner without a browser login
On CI, skip sume login: install the CLI, run sume auth setup with an API key from a secret, and confirm with sume auth status before any job step.
- sume/auto for a former Sora feature: when to pin a model
Sume's sume/auto picks a family and never says which. Good for general clips, wrong when a brand needs one look. How to choose between auto and a pinned id.
- sume/auto for images: no model named, no seed, so pin ids for brand
sume/auto picks an image family and never says which, and there is no seed. When Auto is fine, and when to pin an id like GPT Image 2.5.
Written by Sume