Failed Sume job over MCP: read jobs_get, because jobs_result gives 409

For a failed Sume job, jobs_result returns 409 job_not_completed. Read error.public_reason with jobs_get, report it, and do not resubmit the same paid payload.

5 min readSume
All posts

When a Sume job fails, read it with jobs_get, not jobs_result. The result call is only for completed jobs, and on a failed job it answers 409 job_not_completed. The failure reason lives on the job record in error.public_reason. The hosted MCP server builds this advice into the tool answer itself: for a terminal failure it tells the agent to read error.public_reason with jobs_get, to report it in the user's language, and not to resubmit the identical payload or spend again without asking.

An agent that gets the order wrong has two bad habits available. It can loop on jobs_result and keep seeing a 409, or it can submit the same paid create again and pay twice for the same failure.

The same rule explains a common transcript: an agent finishes a wait, sees failed, calls jobs_result, gets a 409, and then decides that the system is broken. Nothing is broken. The 409 is the API saying that there is no result to return for this job, and the answer to the real question is somewhere else.

Which call for which status

Sume jobs have five statuses. queued and processing are working states; completed, failed and canceled are terminal. Only completed has a result.

A short way to remember the table: results for success, the job record for everything that ended without a result, and waits for anything that is still moving. The canceled row exists because a cancel that arrived before generation began leaves a terminal job with no output, and a second cancel on it just returns the same job.

Status to MCP read tool, from the Sume docs (read 2026-10-05)
StatusTerminalRead withNotes
queuedNojobs_wait or jobs_statusNormal accepted state
processingNojobs_wait or jobs_statusStill billed while it runs
completedYesjobs_resultArtifacts with Sume media URLs
failedYesjobs_getRead error.public_reason
canceledYesjobs_getCancel is idempotent

What the agent envelope tells the model

After a failed answer, the tool result for the agent carries no poll interval. The server sets poll_after_seconds to null on terminal failure, because polling a terminal job cannot change anything. Instead, the next steps in the envelope say to stop and report. A model that follows them will not enter a retry loop.

The envelope also keeps the refund question simple. The agent does not need to reason about billing on a failure; it should say what happened and ask before spending again. If the user agrees, the retry gets a new idempotency_key, because the old key is tied to the failed request.

Using the reason

The public reason is written to be shown to a person. It is not a raw provider code, and public events and results do not expose raw provider task ids or raw provider URLs. That is why an agent can quote it in a message without cleaning it first. A good report has four parts: the job id, the status, the reason in the user's language, and one proposed change, such as a rewritten prompt.

Keep the wording neutral. A reason such as a rejected prompt or an input that could not be fetched is a fact about the request, not a verdict about the person. A good agent offers a rewrite or a different input, and says that nothing else was changed.

When the reason is thin

If the reason is not enough to decide, jobs_events gives the public timeline: job.created, job.queued, job.started, generation.submitted, then job.failed. The stage where the timeline stops tells you whether the failure came before or after generation began.

If every job in a wave failed with the same reason, the problem is probably the shared input, and not one unlucky render. Read each job on its own before you decide.

For a small helper script that classifies the next call, this is enough:

def next_read(status):
    if status == "completed":
        return "jobs_result"
    if status in ("failed", "canceled"):
        return "jobs_get"  # read error.public_reason, then stop
    return "jobs_wait"  # queued or processing: same ids, never resubmit

for s in ("queued", "completed", "failed"):
    print(s, "->", next_read(s))

Retry only on purpose

The key rule is the one on Jobs and results: do not submit the original paid request again only because a read failed or a local process timed out. A retry after a real failure should be a deliberate choice with a changed input and a new idempotency_key; the same key with the same payload returns the original job. The gates for paid calls are in MCP tools and gates.

For batches, remember that a failure on one id does not say anything about the other ids. Read each entry on its own, and use jobs_get only for the ids that did not complete.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume