Failed Sume job over MCP: read jobs_get, because jobs_result gives 409
For a failed Sume job, jobs_result returns 409 job_not_completed. Read error.public_reason with jobs_get, report it, and do not resubmit the same paid payload.

When a Sume job fails, read it with jobs_get, not jobs_result. The result call is only for completed jobs, and on a failed job it answers 409 job_not_completed. The failure reason lives on the job record in error.public_reason. The hosted MCP server builds this advice into the tool answer itself: for a terminal failure it tells the agent to read error.public_reason with jobs_get, to report it in the user's language, and not to resubmit the identical payload or spend again without asking.
An agent that gets the order wrong has two bad habits available. It can loop on jobs_result and keep seeing a 409, or it can submit the same paid create again and pay twice for the same failure.
The same rule explains a common transcript: an agent finishes a wait, sees failed, calls jobs_result, gets a 409, and then decides that the system is broken. Nothing is broken. The 409 is the API saying that there is no result to return for this job, and the answer to the real question is somewhere else.
Which call for which status
Sume jobs have five statuses. queued and processing are working states; completed, failed and canceled are terminal. Only completed has a result.
A short way to remember the table: results for success, the job record for everything that ended without a result, and waits for anything that is still moving. The canceled row exists because a cancel that arrived before generation began leaves a terminal job with no output, and a second cancel on it just returns the same job.
| Status | Terminal | Read with | Notes |
|---|---|---|---|
| queued | No | jobs_wait or jobs_status | Normal accepted state |
| processing | No | jobs_wait or jobs_status | Still billed while it runs |
| completed | Yes | jobs_result | Artifacts with Sume media URLs |
| failed | Yes | jobs_get | Read error.public_reason |
| canceled | Yes | jobs_get | Cancel is idempotent |
What the agent envelope tells the model
After a failed answer, the tool result for the agent carries no poll interval. The server sets poll_after_seconds to null on terminal failure, because polling a terminal job cannot change anything. Instead, the next steps in the envelope say to stop and report. A model that follows them will not enter a retry loop.
The envelope also keeps the refund question simple. The agent does not need to reason about billing on a failure; it should say what happened and ask before spending again. If the user agrees, the retry gets a new idempotency_key, because the old key is tied to the failed request.
Using the reason
The public reason is written to be shown to a person. It is not a raw provider code, and public events and results do not expose raw provider task ids or raw provider URLs. That is why an agent can quote it in a message without cleaning it first. A good report has four parts: the job id, the status, the reason in the user's language, and one proposed change, such as a rewritten prompt.
Keep the wording neutral. A reason such as a rejected prompt or an input that could not be fetched is a fact about the request, not a verdict about the person. A good agent offers a rewrite or a different input, and says that nothing else was changed.
When the reason is thin
If the reason is not enough to decide, jobs_events gives the public timeline: job.created, job.queued, job.started, generation.submitted, then job.failed. The stage where the timeline stops tells you whether the failure came before or after generation began.
If every job in a wave failed with the same reason, the problem is probably the shared input, and not one unlucky render. Read each job on its own before you decide.
For a small helper script that classifies the next call, this is enough:
def next_read(status):
if status == "completed":
return "jobs_result"
if status in ("failed", "canceled"):
return "jobs_get" # read error.public_reason, then stop
return "jobs_wait" # queued or processing: same ids, never resubmit
for s in ("queued", "completed", "failed"):
print(s, "->", next_read(s))Retry only on purpose
The key rule is the one on Jobs and results: do not submit the original paid request again only because a read failed or a local process timed out. A retry after a real failure should be a deliberate choice with a changed input and a new idempotency_key; the same key with the same payload returns the original job. The gates for paid calls are in MCP tools and gates.
For batches, remember that a failure on one id does not say anything about the other ids. Read each entry on its own, and use jobs_get only for the ids that did not complete.
Sources
Related posts
More in Agents
- generate_video MCP defaults to sume/auto; REST requires model
The generate_video MCP tool defaults model to sume/auto, but POST /v1/videos requires model. Same payload schema; the MCP server adds the default.
- jobs_cancel on Sume MCP: Write scope, idempotency_key, early moment
jobs_cancel is a write tool marked destructive. It needs the mcp:write scope or an API key, an idempotency_key, and a job that has not started generating.
- Sume jobs_wait with 20 ids: one unknown id fails the whole call
A batch jobs_wait takes 1 to 20 ids. One unknown or foreign-workspace id makes the full call fail, so validate ids with jobs_list or jobs_status first.
- jobs_wait returned early during a deploy: retry the same job ids
A Sume MCP jobs_wait can return before its 50 second slice when the API host is draining for a deploy. The job is fine: call jobs_wait again on the same ids.
Written by Sume