jobs_wait returned early during a deploy: retry the same job ids

A Sume MCP jobs_wait can return before its 50 second slice when the API host is draining for a deploy. The job is fine: call jobs_wait again on the same ids.

5 min readSume
All posts

A hosted Sume jobs_wait call can come back earlier than its time slice when the API host is draining for a deploy. In the source, the wait loop checks whether the host is draining and stops waiting when it is, and the code checks that flag every second while it sleeps. The answer that you get is the same shape as a normal timed-out wait: the job has not changed state, and you ask again with the same job ids. Do not submit the paid create again and do not call the job failed.

This is a small behavior with a large cost if an agent gets it wrong, because the wrong reaction is a duplicate paid job.

The practical test for your own agent is simple: after any jobs_wait answer that is not terminal, does it call jobs_wait again with the ids it already has? If it instead calls a create tool, it will pay twice. The retry rule is what makes a deploy invisible to the user.

Why the wait ends early

The remote default slice is 50 seconds and the cap is 55. A wait normally returns when the job reaches a terminal state or when the slice ends. The drain check adds a third reason to return early: Sume would rather answer the wait now than hold an HTTP request open while the host shuts down.

Deploys are not the only reason for a short answer. A wait can also end because the caller aborted the request, and the server turns an interrupted wait into a retryable error with a poll_status next action. In both cases the job is not touched: the wait is only a read, and a read cannot change a job.

What the agent should do

The jobs page in the docs gives the rule for a timed-out wait: on wait_slice_expired, retry jobs_wait with the same ids, and never submit the paid create again. An early return during a deploy is handled the same way, so an agent does not need a special branch.

Do not tune the loop to the deploy. Add no extra sleep, no larger timeout, and no special error text for it. A plain retry of the same call is correct, and the server's own guidance in the answer says the same thing.

Reactions to a jobs_wait answer that is not terminal (read 2026-10-05)
What you seeWhat it meansNext call
wait_slice_expiredThe slice ended; the job still runsjobs_wait with the same ids
Early return, job not terminalHost drained or slice endedjobs_wait with the same ids
524, 522, 523 or 525Transport failure at the edgejobs_wait again, or one jobs_status read
operator_stoppedSume operations stopped an idStop; the answer will not change

Keep the ids and keep going

Keep the ids in the agent's state, because that is all the retry needs. After a few rounds without progress, do one jobs_status read and look at the status before you decide anything. A job is in queued, processing, completed, failed or canceled; only the last three are terminal.

A tiny loop in the shape that the docs describe, written for a script that talks to the HTTP API with a key, is a good reference for how little state is needed:

If you run your own wrapper around the MCP client, count the rounds and log the ids, not the prompts. When a render takes ten minutes, you will see about a dozen slices, and some of them will be short. That is normal.

import time

def wait_for(job_ids, call_wait):
    # call_wait(ids) returns a dict with a terminal flag
    while True:
        answer = call_wait(job_ids)
        if answer.get("terminal"):
            return answer
        time.sleep(1)  # same ids again; never resubmit the paid create

Same rule as every long render

The longer rule is the same as the one that governs a ten-minute render: do not ask for a longer wait, ask again. Each slice is one HTTP request, and edges close idle requests, so the design is many short waits. See Jobs and results for the status table, and MCP tools and gates for the tool inventory.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume