jobs_wait wait_for any: other Sume jobs keep running and billing

jobs_wait takes up to 20 job_ids and wait_for all or any. With any it returns on the first finish, but the other jobs keep running and billing.

5 min readSume
All posts

On Sume's hosted MCP, jobs_wait accepts either one job_id or a job_ids array of 1 to 20, and wait_for is all or any. With any the call can return as soon as one job finishes, but it still reports every id, and the remaining jobs keep running and keep billing. Choosing any changes when you get control back, not what you pay.

This comes from the jobs_wait tool contract in the Sume MCP server and from MCP tools and gates, read on 2026-10-03.

The call shape

Pass exactly one of job_id or job_ids. The optional fields are wait_for, timeout_seconds, interval_seconds and include_results. On remote MCP every caller uses timeout_seconds of 45 to 55 and re-issues the same call until the jobs are terminal; a longer value is clamped to 55 seconds, and there is no server-side hold beyond that. The default interval is 5 seconds and the maximum is 60.

jobs_wait parameters on remote MCP (read 2026-10-03)
ParameterValueNote
job_id or job_ids1 to 20 idsExactly one of the two
wait_forall or anyany still reports every id
timeout_seconds45 to 55Longer values clamp to 55
interval_secondsDefault 5, max 60Never busy-loop under 5 seconds
include_resultstrue or falseReturns finished results inline

When any is useful, and its cost

any suits a race where you want the first usable clip and plan to ignore the rest, or an agent that wants to start inspecting one result while the others render. It does not cancel anything. The other jobs are billed and still running, so if you do not want them, cancel them explicitly with jobs_cancel, which is a write tool and needs an idempotency_key.

Abandoning jobs because one jobs_wait call returned is the expensive mistake: it orphans billed, still-running work that nobody will collect.

Reading the outcome

The response says which clock ended the wait. wait_slice_expired means your polling window closed, not that the jobs failed; pending jobs are still running and you should re-issue the call on the pending_job_ids it lists. operator_stopped means operations stopped those ids: it is terminal, there is no output, and you should stop polling them.

A 429 rate_limited on the wait is a throttled read, never a job outcome. Back off for retry_after_seconds and call again. Reads and submits have separate budgets, so a throttled read does not refuse your creates.

Sizing the wave

The 20-id limit caps a single wait call, not the wave. For larger sets, split the ids across several calls. Capacity is a separate question answered by generation admission, which reports queue capacity remaining, so size a wave to that number rather than to 20.

A pattern that does not leak spend

Keep a set of pending ids in the agent's working notes. After each jobs_wait return, move finished ids to a done set, then decide for each pending id: keep waiting, or cancel. Re-issue the wait with the same ids until the pending set is empty or you have cancelled what remains.

Use include_results: true when you want the finished results inline, so no separate jobs_result call is needed. Do this when the result is small; for large outputs fetch them one at a time.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume