Eight concurrent jobs_wait calls per principal on Sume hosted MCP
Sume's hosted MCP bounds waits per process and per principal, workspace plus owner. Batch ids into one jobs_wait instead of sending many parallel waits.

Sume's hosted MCP server keeps operational budgets on how many requests it handles at once. Per process, Sume's operations notes list 20 waits, with eight per principal, where a principal is a workspace plus an owner. The practical rule is simple: send one batch jobs_wait for up to 20 ids, not 20 single waits in parallel. These are operational limits and can change, so read them as a design hint, not a contract.
The budgets
A principal is workspace plus owner, so two different owners in one workspace do not share one principal's count. The numbers apply to one API process, not to your whole account.
| Request kind | Per process | Per principal |
|---|---|---|
| Reads | 32 | 4 |
| Waits (jobs_wait) | 20 | 8 |
| Paid creates | 64 | 20 |
| Other tool requests | 16 | 4 |
What happens at the limit
Reads and waits are answered at once instead of queuing. A saturated wait returns the usual wait_slice_expired result with all ids kept. Other tools return a retryable wait_busy error with HTTP 429 and a one-second retry hint. Paid creates and other tools may wait up to 20 seconds for a slot before they are refused.
A refused call creates no job and spends nothing, and idempotency and wallet holds are unchanged. So the safe move is the same as for any transport error: retry that call with the same idempotency_key.
- Batch ids: one
jobs_waittakes up to 20 ids, andjobs_resulttakes the same 20. - Do not fan out a wait per job from several sub-agents under one key.
- Back off after a busy answer rather than retrying in a tight loop.
The tradeoff
Batching means a single slow job in the group holds the whole answer until the slice ends, unless you set wait_for to any. With any, the call returns early, but the remaining jobs keep running and billing, so you must call again for them. A batch of similar jobs is the common case where all is right.
If you have many agents sharing one workspace and owner, they share these counts. That is a reason to put a single waiter in front of them, not to raise parallelism.
Sources
Related posts
More in Developers
- Envoy's 15 s route timeout vs a 30 s Sume sync call
Envoy's RouteAction timeout defaults to 15s. Sume sync mode can wait 30 s. Behind Envoy or Istio, use async, or raise the route timeout above the sync wait.
- Estimate a batch image cost from the Sume catalog in Python
Read a model's price from GET /v1/images/models/{id}/endpoints and multiply by n and count. A 14-line Python estimator, with the quality and size caveat.
- Excel VBA: submit an AI video job and poll it with Application.OnTime
An Excel macro posts a Wan 3.0 job to Sume with ServerXMLHTTP, saves the job id in a cell and re-checks status with Application.OnTime without blocking Excel.
- Express breaks your Sume webhook signature: mount express.raw first
If verifyWebhook returns false in Express, a JSON parser changed the body. Mount express.raw on the webhook route only. Working Node code with the SDK.
Written by Sume