jobs_events after jobs_get: which stage did a failed Sume job stop at?

Over Sume MCP, jobs_get gives a failed job's public_reason and retryable flag; jobs_events gives the stage timeline. When to call each.

5 min readSume
All posts

When a paid job fails over Sume's hosted MCP, read jobs_get first for the reason and the retryable flag, then jobs_events for the stage the job stopped at. jobs_result is the wrong tool here: on a failed job it answers a 409 with no error fields. And a failed job is not a signal to resubmit the paid create; look at retryable before spending again.

The tool guidance below comes from the Sume MCP tool contracts, plus MCP tools and gates and Jobs and results, read on 2026-10-03.

Three reads, three jobs

jobs_get returns the job record, and it is the only tool that can tell you why a job failed. On a failed job it carries the whole error: error.public_reason, a sanitized error.message, and error.retryable. The tool contract points out that error.code is shared by nearly every generation failure, so branch on public_reason, not on the code.

jobs_events lists sanitized lifecycle events for one job, with a limit that defaults to 50. It is for stalls, failures and stage transitions after status looks odd. It is not for routine polling and it does not return output.

Which job read to use when (read 2026-10-03)
SituationToolWhy
Waiting for a resultjobs_waitPolls in slices up to 55 seconds
One quick status checkjobs_statusReturns immediately
Job succeededjobs_resultReturns artifacts
Job failed, why?jobs_getpublic_reason, message, retryable
Job failed or stalled, where?jobs_eventsStage transitions

The 409 trap

jobs_result on a job that is still queued or processing returns 409 job_not_completed. That is not a failure; the job is billed and still running, and the right move is to keep waiting on the same job id. But the same 409 also comes back for a job that has already failed, and in that case it carries no error fields. If an agent reads every 409 as a reason to resubmit, it will pay twice for work that failed for a reason it never looked up.

The safe rule for an agent: on a 409 from jobs_result, call jobs_status once. If the status is non-terminal, go back to jobs_wait. If it is failed, call jobs_get, then jobs_events.

Deciding whether to retry

retryable is the first gate. If it is false, a second submit with a new idempotency_key will most likely fail the same way, and the message usually names what to change: a duration, a reference image, or a model limit. Fix the input, then submit once.

If it is true, retry with care. Reusing the same idempotency_key returns the original job rather than starting a new one, which is what you want when you are unsure whether the first submit landed. A new key starts new billable work. Send a new key only when you have decided the first job is dead and the input is correct.

Keep the output safe

Job reads can include signed or private media URLs. Tool guidance is to keep them out of user-visible reports and logs. When an agent summarizes a failure, quote public_reason and the stage from jobs_events, and leave the URLs out.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume