Sume MCP jobs_wait returns 524: it is not a failed job
A 524, 522, 523 or 525 on Sume jobs_wait is a transport failure, not a job result. Wait again on the same ids and never resubmit the paid create.

A 524 on jobs_wait means the connection was cut while the request sat open. It says nothing about the job. The job keeps running and billing, so call jobs_wait again with the same ids, or read jobs_status once. Never submit the paid create again and never report the job as blocked.
| What you see | Meaning | Action |
|---|---|---|
wait_slice_expired | Slice ended, job not terminal | Wait again, same ids |
wait_slice_clamped | Requested timeout above the cap | Use 55 s or less |
| 524, 522, 523, 525 | Transport failure | Wait again or read jobs_status |
502 / Transport send error | Request held too long at the edge | Use slices of 55 s or less |
operator_stopped | Sume operations stopped an id | Terminal, holds refunded |
Why it happens
A wait is one HTTP request held open for the whole slice, with no data moving. Edges close such requests at some time. The Sume docs say the old thread-header 600-second server hold is gone, and a request open that long would die at the edge with a 502 before it could answer.
The safe loop
The danger is the retry that creates a second paid job. Keep the create and the wait separate in the agent's plan.
- Create once. Save the returned job id and the
idempotency_keyyou sent. - Wait with the default slice (50 s).
- On any 52x or
wait_slice_expired, wait again on the same id. - If the wait keeps failing, read
jobs_statusonce to see the state. - Read
jobs_resultonly after the job is terminal.
Batch waits and partial answers
With job_ids you can wait on up to 20 jobs. wait_for: any returns when one is done, but the others continue and still bill. An unknown id, or an id from another workspace, fails the whole call, so check ids first when a batch wait errors at once.
What to log
Log the job id, the number of waits, and each status code. When a wait returns 524, record it but do not alarm: it is expected on long holds. If the same id fails the wait many times in a row, switch to jobs_status with backoff, which is a quick read and does not hold a connection.
Reporting to the user
Say what is true: the wait connection dropped, the job is still being processed, and you are waiting again. Do not say the job failed unless jobs_status says failed, and do not offer to resubmit while the job is processing.
Some teams add a hard ceiling on total waiting, for example thirty minutes per job, after which the agent reports the state and asks. That is reasonable, and it is better than resubmitting. The ceiling should apply to the wait loop, not to the job: the job may still finish later, and the agent can read the result with jobs_result when it does.
Before you rely on this setup, run a short acceptance test with a read-only credential. Connect, call mcp_health, call tools_list, and read one job with jobs_status. Record the tool count you see, so you can notice later if a credential change alters it. Then repeat the test after any config edit. A five-minute test like this catches most wiring mistakes before they cost money, and it gives you a baseline to compare against when something behaves differently next week.
Sources
More in Integrations
- Sume MCP jobs_wait with 20 job_ids: wait_for any versus all
jobs_wait takes up to 20 job_ids and wait_for any or all. All is the default. Any returns early, but the other jobs keep running and billing. Examples inside.
- Sume MCP with Write off: what an OAuth agent can still do
With OAuth read-only (mcp:read), an agent on Sume's hosted MCP can still read jobs, assets and the catalog and scrape pages. Writes and paid tools fail.
- Sume 429 queue_full from an MCP agent: wave sizes for Free to Scale
An agent that submits too many paid jobs hits 429 queue_full. Accepted capacity is 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Wave sizes inside.
- Sume MCP script_run: fan out 3+ calls in one turn, 55 s limit
script_run runs a short JavaScript program on Sume that calls tools in a loop or in parallel. timeout_seconds is 5 to 55. Use it for 3+ similar calls.
Written by Sume