Pro key: 300 writes and 12,000 reads a minute for an agent on Sume MCP

On Sume Pro, each key gets 300 writes and 12,000 reads per minute. An agent polling 20 jobs every 5 s uses 240 reads, 2 percent of the read budget.

4 min readSume
All posts

A Sume Pro key may make 300 write requests and 12,000 read requests per minute. An agent that polls 20 jobs every 5 seconds makes 240 reads a minute, which is 2 percent of the read budget, so polling will not hit a 429 before job capacity does. Sume sized reads at forty times the write number because agents poll many jobs, and reads and writes have separate budgets.

The table

From Sume's Authentication page (read 2026-10-08): the plan of the workspace that owns the key sets the budget for all of /v1. A read is any GET or HEAD, and also the two POSTs that submit nothing: /v1/generation/admission-preview and the MCP endpoint itself.

API key request budgets per minute, from Sume docs read 2026-10-08
PlanWritesReadsReads / writes
Free1204,80040
Pro30012,00040
Startup60024,00040
Scale1,20048,00040

What an MCP agent spends

An MCP tool call spends the write budget once, for the run it creates. It does not spend write budget for the JSON-RPC request that carried it. A jobs_status poll over MCP spends no write budget at all. So a loop that submits one clip and waits uses one write and then reads.

Take twenty running jobs, polled with jobs_status every 5 seconds: 20 x (60 / 5) = 240 reads per minute. On Free that is 240 / 4,800 = 5 percent, on Pro 240 / 12,000 = 2 percent. A batch jobs_wait on 20 ids needs far fewer calls, about one per 55 seconds.

Which limit an agent hits first

Rate is not capacity. Concurrency and queue size cap how many generations run, and they do not grow with request rate. A Pro workspace accepts 24 paid jobs (4 processing plus 20 queued), so the 25th concurrent submit returns 429 queue_full long before it nears 300 writes a minute. Read the 429 body: error.details.scope names read or write when it is a rate-limit error, and queue_full is a separate code.

  • rate_limited: back off, use retry-after when present.
  • queue_full: wait for jobs to finish or cancel queued ones, then retry with the same idempotency key.
  • Read ratelimit-remaining on responses instead of counting requests yourself.

Practical settings for the agent

Tell the agent to prefer one batch jobs_wait over many jobs_status polls, to use at most 20 ids per call, and to honor retry-after on a 429. Unauthenticated requests are limited per client IP at the Free rate, with a read bucket of four times the writes, so an anonymous probe does not get the agent-sized budget. The multiple is a deployment configuration value, so ratelimit-limit on the response is the authority for the host you use.

A quick budget check

Before you ship an agent, multiply. Writes: how many paid or mutating calls per minute at peak? Reads: jobs times polls per minute, plus any list calls. If the writes exceed 300 on Pro, move to Startup (600) or Scale (1,200), or spread the work over more minutes. If reads exceed 12,000, switch from per-job polling to batch waits, which is the cheaper fix. Both numbers are per key budgets for the workspace plan, as read 2026-10-08.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume