Eight concurrent jobs_wait calls per principal on Sume hosted MCP

Sume's hosted MCP bounds waits per process and per principal, workspace plus owner. Batch ids into one jobs_wait instead of sending many parallel waits.

4 min readSume
All posts

Sume's hosted MCP server keeps operational budgets on how many requests it handles at once. Per process, Sume's operations notes list 20 waits, with eight per principal, where a principal is a workspace plus an owner. The practical rule is simple: send one batch jobs_wait for up to 20 ids, not 20 single waits in parallel. These are operational limits and can change, so read them as a design hint, not a contract.

The budgets

A principal is workspace plus owner, so two different owners in one workspace do not share one principal's count. The numbers apply to one API process, not to your whole account.

Per-process outer budgets for hosted MCP (Sume operations notes, read 2026-10-06; may change)
Request kindPer processPer principal
Reads324
Waits (jobs_wait)208
Paid creates6420
Other tool requests164

What happens at the limit

Reads and waits are answered at once instead of queuing. A saturated wait returns the usual wait_slice_expired result with all ids kept. Other tools return a retryable wait_busy error with HTTP 429 and a one-second retry hint. Paid creates and other tools may wait up to 20 seconds for a slot before they are refused.

A refused call creates no job and spends nothing, and idempotency and wallet holds are unchanged. So the safe move is the same as for any transport error: retry that call with the same idempotency_key.

  • Batch ids: one jobs_wait takes up to 20 ids, and jobs_result takes the same 20.
  • Do not fan out a wait per job from several sub-agents under one key.
  • Back off after a busy answer rather than retrying in a tight loop.

The tradeoff

Batching means a single slow job in the group holds the whole answer until the slice ends, unless you set wait_for to any. With any, the call returns early, but the remaining jobs keep running and billing, so you must call again for them. A batch of similar jobs is the common case where all is right.

If you have many agents sharing one workspace and owner, they share these counts. That is a reason to put a single waiter in front of them, not to raise parallelism.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume