Sume MCP jobs_wait returns 524: it is not a failed job

A 524, 522, 523 or 525 on Sume jobs_wait is a transport failure, not a job result. Wait again on the same ids and never resubmit the paid create.

5 min readSume
All posts

A 524 on jobs_wait means the connection was cut while the request sat open. It says nothing about the job. The job keeps running and billing, so call jobs_wait again with the same ids, or read jobs_status once. Never submit the paid create again and never report the job as blocked.

Wait outcomes and what to do, read 2026-10-08
What you seeMeaningAction
wait_slice_expiredSlice ended, job not terminalWait again, same ids
wait_slice_clampedRequested timeout above the capUse 55 s or less
524, 522, 523, 525Transport failureWait again or read jobs_status
502 / Transport send errorRequest held too long at the edgeUse slices of 55 s or less
operator_stoppedSume operations stopped an idTerminal, holds refunded

Why it happens

A wait is one HTTP request held open for the whole slice, with no data moving. Edges close such requests at some time. The Sume docs say the old thread-header 600-second server hold is gone, and a request open that long would die at the edge with a 502 before it could answer.

The safe loop

The danger is the retry that creates a second paid job. Keep the create and the wait separate in the agent's plan.

  • Create once. Save the returned job id and the idempotency_key you sent.
  • Wait with the default slice (50 s).
  • On any 52x or wait_slice_expired, wait again on the same id.
  • If the wait keeps failing, read jobs_status once to see the state.
  • Read jobs_result only after the job is terminal.

Batch waits and partial answers

With job_ids you can wait on up to 20 jobs. wait_for: any returns when one is done, but the others continue and still bill. An unknown id, or an id from another workspace, fails the whole call, so check ids first when a batch wait errors at once.

What to log

Log the job id, the number of waits, and each status code. When a wait returns 524, record it but do not alarm: it is expected on long holds. If the same id fails the wait many times in a row, switch to jobs_status with backoff, which is a quick read and does not hold a connection.

Reporting to the user

Say what is true: the wait connection dropped, the job is still being processed, and you are waiting again. Do not say the job failed unless jobs_status says failed, and do not offer to resubmit while the job is processing.

Some teams add a hard ceiling on total waiting, for example thirty minutes per job, after which the agent reports the state and asks. That is reasonable, and it is better than resubmitting. The ceiling should apply to the wait loop, not to the job: the job may still finish later, and the agent can read the result with jobs_result when it does.

Before you rely on this setup, run a short acceptance test with a read-only credential. Connect, call mcp_health, call tools_list, and read one job with jobs_status. Record the tool count you see, so you can notice later if a credential change alters it. Then repeat the test after any config edit. A five-minute test like this catches most wiring mistakes before they cost money, and it gives you a baseline to compare against when something behaves differently next week.

Sources

More in Integrations

All Integrations posts

Written by Sume