Enterprise API rate limit: what a Sume key gets by default
Sume's Enterprise rate limit is contracted, not self-serve. Until a number is provisioned, an Enterprise key gets the Scale row, plus 20 concurrent jobs.

Enterprise rate limits on Sume are set by contract, not self-serve: the Authentication page lists "Contact sales" for both the write and read budgets. Until a contracted number is provisioned, an Enterprise key resolves to the Scale row: 1200 writes and 48000 reads per minute. Generation concurrency is a separate limit, and its Enterprise default is 20.
This comes from Authentication and Generation admission, read 2026-09-29. Both pages say to trust the values on a live response over the tables.
What are the request limits by plan?
Every API key gets a per-minute request budget across all of /v1, set by the subscription plan of the workspace the key belongs to. Reads and writes have separate budgets, and the read number is forty times the write number.
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
| Enterprise | Contact sales | Contact sales |
What does an Enterprise key get before a contract?
The page says Enterprise is not self-serve, and that until a contracted number is provisioned an Enterprise key resolves to the Scale row above. So the starting point is 1200 writes and 48000 reads a minute. A contract changes that number; nothing in the docs gives you a way to raise it yourself.
How many jobs can an Enterprise workspace run at once?
That is concurrency, not request rate. The docs say raising your request rate does not raise it. Generation concurrency counts paid jobs in processing; extra valid jobs wait as queued while queue capacity remains.
| Plan | Processing | Queue capacity (default) | Accepted jobs |
|---|---|---|---|
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
Where do I read my real limits?
Two places. On any response, ratelimit-limit is the authority for the deployment you are talking to, and ratelimit-remaining and ratelimit-reset describe the current window. For concurrency, the dashboard Concurrency tab is the source of truth, exposed as generation_limits.concurrency_limit. When an admin override raises it, limit_source reads admin_override.
curl -sS -D - -o /dev/null https://api.sume.com/v1/me \
-H "Authorization: Bearer $SUME_API_KEY" \
| grep -i '^ratelimit'Can I buy more concurrency with a top-up?
No. Generation concurrency is plan-only, and prepaid top-ups do not raise it. For Enterprise, the docs say higher contract limits go through admin overrides. They also say organization workspaces have a floor of 10.
When you hit a limit, tell them apart: a 429 rate_limited names read or write in error.details.scope, while 429 queue_full means concurrency plus queue capacity is full (queueing).
What should I ask for in a contract?
The docs give you three numbers that a contract can move, so ask about each by name.
- The write budget per minute. This is the number the plan buys; reads are derived from it.
- The read budget per minute, if you poll many jobs at once. Agent-style harvests can hold twenty-odd jobs open and poll each one.
- The concurrency limit, and with it the queue capacity. Concurrency is what decides how many paid jobs run at once, and queue capacity defaults to the larger of 3 and five times the concurrency limit.
Sources
Related posts
More in Developers
- Export API usage to CSV: turning Sume's usage ledger into rows
Sume's docs list no CSV export, but GET /v1/usage returns ledger rows as JSON. Convert them with jq, and keep captured rows separate from refunds.
- fal.ai API rate limit: concurrency from 2 up to 40
fal.ai limits how many requests run at once, not requests per minute: 2 for a new account, rising with credit purchases to 40 self-serve.
- fal AI FFmpeg API: endpoints, inputs and prices
fal hosts FFmpeg as model endpoints: merge videos, merge audio and video, extract a frame, compose tracks. Called with a fal key, priced per second.
- fal bytedance/seedance-2.0 request fields, mapped to Sume's
Moving a fal Seedance 2.0 call to Sume: the path becomes a bare model id, image_urls become input_references, and seed, auto ratio and auto duration go.
Written by Sume