Enterprise API key rate limit before a contract: the Scale row

Until Sume provisions a contracted number, an Enterprise key uses the Scale row: 1,200 writes and 48,000 reads per minute. The full plan table and the math.

4 min readSume
All posts

Until Sume provisions a contracted number, an Enterprise API key uses the Scale row of the rate-limit table: 1,200 writes per minute and 48,000 reads per minute. The Authentication page says Enterprise is not self-serve, and that the table's Enterprise cell reads "Contact sales" until a number is agreed.

The table

Each API key gets a per-minute request budget for all of /v1, set by the subscription plan of the workspace that owns the key. Reads and writes have separate budgets.

Per-key request budgets per minute (Sume docs, read 2026-10-09)
PlanWrites per minuteReads per minuteReads divided by writes
Free1204,80040
Pro30012,00040
Startup60024,00040
Scale1,20048,00040
EnterpriseContact sales (Scale row until provisioned)Contact sales (Scale row until provisioned)40

What counts as a write

A read is any GET or HEAD: a poll of status_url, events_url, or result_url, or a list call. Two POST routes that submit nothing also count as reads: /v1/generation/admission-preview and the MCP endpoint itself. Everything else is a write, including run creation, cancellation, and uploads.

Because the buckets are separate, a tight status-poll loop cannot cause a 429 on your own submits. A 429 names the exhausted budget in error.details.scope, either read or write.

Request rate is not generation capacity

The request budget is not the number of jobs that run at once. Concurrency comes from the plan and is reported in generation_limits. The docs state that raising your request rate does not raise the concurrency limit. For Enterprise, the default processing concurrency is 20 and the default accepted capacity is 120, with admin overrides for higher contract limits.

So a Scale-row Enterprise key can send 1,200 writes a minute, but at the default only 120 paid generation jobs can be processing or queued at once. The rest get 429 queue_full. Read the two limits separately in your client.

Do not count requests yourself

The docs ask you to read ratelimit-remaining instead of keeping your own counter, and to wait for retry-after after a 429. The headers describe the budget that the current request spent from. The ratelimit-limit header is the authority for the deployment you call, because the read multiple is a deployment configuration value. See the curl header post for a one-line check.

curl -s -D - -o /dev/null https://api.sume.com/v1/me \
  -H "Authorization: Bearer $SUME_API_KEY" | grep -i '^ratelimit'

Planning around the shared row

Because Enterprise reuses the Scale row until a contract sets different limits, plan for 1,200 writes and 48,000 reads per key per minute. Reads and writes sit in separate buckets, so a busy poller does not use up your submit budget.

Each key has its own bucket. Splitting traffic across keys gives each key its own allowance, but a 429 carries error.details.scope so you can see which bucket you hit. Back off on retry-after rather than guessing a delay, and log ratelimit-remaining so you see the limit approaching before the first 429.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume