Sume API rate limits by plan: requests per minute for writes and reads

Sume gives every API key a per-minute budget set by plan: 120 writes on Free up to 1200 on Scale, with reads at forty times the write number. Table and headers.

5 min readSume
All posts

The Sume API gives each key a request budget per minute across all of /v1, set by the plan of the workspace the key belongs to: 120 writes a minute on Free, 300 on Pro, 600 on Startup and 1200 on Scale. Reads have their own bucket at forty times the write number, so a tight status-poll loop cannot 429 your own submits.

The numbers

A read is any GET or HEAD, plus two POSTs that submit nothing. Everything else, including creating runs, cancelling and uploads, is a write. Enterprise is contract-based; until a number is provisioned an Enterprise key resolves to the Scale row.

Sume API request budget per key per minute (read 2026-10-03)
PlanWrites per minuteReads per minute
Free1204800
Pro30012000
Startup60024000
Scale120048000
EnterpriseContact salesContact sales

Do not confuse it with generation capacity

Request rate is a different control from how many generations run at once. Concurrency and queue capacity come from your plan's generation_limits, and raising your request rate does not raise them. A 429 can mean either: rate_limited is the request budget, queue_full is the generation queue.

Read the headers instead of counting

Responses carry ratelimit-limit, ratelimit-remaining, ratelimit-reset and, on a 429, retry-after. They describe whichever bucket the current request spent from, and a 429 names it in error.details.scope as read or write. The read multiple is a deployment setting, so the headers on the response are the authority for the host you call, not this table.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume