MCP server rate limit: per tool call or per request on Sume?
On Sume the MCP endpoint POST is a read, a tool call spends write budget once for the run it creates, and a jobs_status poll over MCP spends no write budget.

On Sume the rate limit follows the work, not the JSON-RPC envelope. The MCP endpoint itself is counted as a read, a tool call spends the write budget once for the run it creates, and a jobs_status poll over MCP spends no write budget at all.
What do the per-minute budgets look like?
Every API key gets a request budget per minute across all of /v1, set by the workspace's plan. Reads and writes have separate budgets, and the plan number is the write number; reads get forty times that.
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
Which budget does the MCP request spend?
A read is any GET or HEAD, plus the two POSTs that submit nothing: /v1/generation/admission-preview and the MCP endpoint itself. So the request that carries a tool call is a read. The write budget is spent by what the tool does: an MCP tool call spends it once, for the run it creates, not for the JSON-RPC request that carried it.
Does polling over MCP cost anything?
A jobs_status poll over MCP spends no write budget at all. Reads are deliberately cheap, sized for an agent that holds twenty-odd jobs open and polls each, so a tight status loop cannot 429 your own submits.
What still limits a busy agent?
Request rate is not generation capacity. How many generations run at once is governed by your plan's concurrency limit, and raising your request rate does not raise it. Read ratelimit-remaining instead of counting requests, and on a 429 look at error.details.scope, which is read or write. Queue capacity is a third control, covered in the generation admission docs.
How many run-creating tool calls fit in a minute?
Divide the write column by the calls you make. A Pro workspace has 300 writes per minute, so up to 300 run-creating tool calls, and a Free workspace has 120. The status polling those runs need comes out of the separate, forty-times-larger read bucket. Enterprise is not self-serve: until a contracted number is provisioned, an Enterprise key resolves to the Scale row.
What does this mean for a gateway in front of MCP?
A gateway that meters requests should classify the MCP POST as a read and not multiply it by the tools inside. The budget that matters for spend is the write bucket, which Sume charges once per run created. Back off on retry-after when a 429 arrives, and use the scope in the error to know which bucket to slow down.
Sources
Related posts
More in Developers
- MCP-Method header: route and rate-limit MCP requests at a gateway
The 2026-07-28 MCP spec requires Mcp-Method and Mcp-Name headers on HTTP requests. How Sume counts MCP calls against its rate limits.
- MCP input_required, inputResponses and multi round trip vs dry_run
MCP 2026-07-28 lets a server return input_required and the client retry with inputResponses. Sume's paid confirmation is two calls with dry_run instead.
- MCP OAuth resource indicator: is a token bound to the server?
Sume's MCP OAuth resource audience is https://mcp.sume.com/mcp, its authorization server is the MCP origin, and the token is not an API key.
- MCP roots, sampling and logging deprecated: Sume impact
The 2026-07-28 MCP spec deprecates Roots, Sampling and Logging. Sume's hosted server declares only tools, so a client calling it has nothing to migrate.
Written by Sume