60 clips on hosted MCP: three jobs_wait batches of 20, 3 calls total

jobs_wait takes 1 to 20 job_ids. For 60 clips that is three batch waits instead of 60. Add include_results to skip the separate result reads too.

3 min readSume
All posts

Why batch

After a parallel fan-out on hosted MCP, prefer one batch jobs_wait to many single waits. jobs_wait accepts a single job_id, or job_ids with 1 to 20 ids and an optional wait_for of all (the default) or any. The batch response is object: "job_wait_batch" with a status snapshot for each id.

Sixty clips is three groups of 20. The call count comparison is simple arithmetic from that ceiling.

Tool calls to wait on and read 60 jobs; arithmetic from the 1-20 id ceiling, read 2026-10-08
ApproachWait callsResult callsTotal
Single wait, single read per job6060120
Batch wait of 20, batch jobs_result of 20336
Batch wait of 20 with include_results30 to 33 to 6

The request

Each call still holds at most 55 seconds, with a default of 50. When it runs out, the outcome is wait_slice_expired. Retry the same ids. The remaining jobs continue and still bill, so never resubmit the paid creates. With wait_for: "any", the response still reports every id.

{
  "job_ids": ["job_a", "job_b", "job_c"],
  "wait_for": "all",
  "include_results": true,
  "timeout_seconds": 55
}

include_results and what it leaves out

With include_results: true, each completed id comes back with its jobs_result answer in results[], so the wave needs no separate read. Some results do not fit in one answer. results_omitted.job_ids names those, and you read them with one batch jobs_result.

jobs_result also takes job_ids with the same 1 to 20 ceiling. The response is job_result_batch, with results[] in request order. Each entry has ok and either a value or a typed error. Partial success is normal: an id still running comes back job_not_completed, and partial_failure.failed_job_ids names exactly the ids that need a second read.

Mistakes that cost money

Unknown ids and ids from another workspace make the whole call fail, so a typo in one id of twenty fails all twenty. Validate ids before you batch them.

Do not treat wait_slice_expired as a failure of the job, and do not call the create tool again to recover. Store the job ids from the creates, because they are the only way to rejoin the work after a restart.

  • Group ids in 20s. Two groups of 30 are not allowed.
  • Keep each paid create's idempotency_key stable on retry, so a repeated create returns the original job.
  • Report media.sume.com ids and URLs in logs, not signed URLs or tokens.

Order of operations for a wave

Create all 60 jobs first, each with its own stable idempotency_key, and write every returned job id to your own store. Then split the ids into three lists of 20 and wait on each list in turn. Because each wait returns when its slice ends, the second and third waits often return at once, since those jobs finished while you were waiting on the first.

When a list comes back with some ids still running, wait on just those ids. Shrinking the list keeps responses small and keeps results_omitted rare.

Logging the wave

Log one line per batch: the number of ids, the wait_for value, how many came back terminal, and the ids named in results_omitted. Do not log signed URLs or tokens. That is enough to explain a slow wave later without storing private data.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume