jobs_events after jobs_get: which stage did a failed Sume job stop at?
Over Sume MCP, jobs_get gives a failed job's public_reason and retryable flag; jobs_events gives the stage timeline. When to call each.

When a paid job fails over Sume's hosted MCP, read jobs_get first for the reason and the retryable flag, then jobs_events for the stage the job stopped at. jobs_result is the wrong tool here: on a failed job it answers a 409 with no error fields. And a failed job is not a signal to resubmit the paid create; look at retryable before spending again.
The tool guidance below comes from the Sume MCP tool contracts, plus MCP tools and gates and Jobs and results, read on 2026-10-03.
Three reads, three jobs
jobs_get returns the job record, and it is the only tool that can tell you why a job failed. On a failed job it carries the whole error: error.public_reason, a sanitized error.message, and error.retryable. The tool contract points out that error.code is shared by nearly every generation failure, so branch on public_reason, not on the code.
jobs_events lists sanitized lifecycle events for one job, with a limit that defaults to 50. It is for stalls, failures and stage transitions after status looks odd. It is not for routine polling and it does not return output.
| Situation | Tool | Why |
|---|---|---|
| Waiting for a result | jobs_wait | Polls in slices up to 55 seconds |
| One quick status check | jobs_status | Returns immediately |
| Job succeeded | jobs_result | Returns artifacts |
| Job failed, why? | jobs_get | public_reason, message, retryable |
| Job failed or stalled, where? | jobs_events | Stage transitions |
The 409 trap
jobs_result on a job that is still queued or processing returns 409 job_not_completed. That is not a failure; the job is billed and still running, and the right move is to keep waiting on the same job id. But the same 409 also comes back for a job that has already failed, and in that case it carries no error fields. If an agent reads every 409 as a reason to resubmit, it will pay twice for work that failed for a reason it never looked up.
The safe rule for an agent: on a 409 from jobs_result, call jobs_status once. If the status is non-terminal, go back to jobs_wait. If it is failed, call jobs_get, then jobs_events.
Deciding whether to retry
retryable is the first gate. If it is false, a second submit with a new idempotency_key will most likely fail the same way, and the message usually names what to change: a duration, a reference image, or a model limit. Fix the input, then submit once.
If it is true, retry with care. Reusing the same idempotency_key returns the original job rather than starting a new one, which is what you want when you are unsure whether the first submit landed. A new key starts new billable work. Send a new key only when you have decided the first job is dead and the input is correct.
Keep the output safe
Job reads can include signed or private media URLs. Tool guidance is to keep them out of user-visible reports and logs. When an agent summarizes a failure, quote public_reason and the stage from jobs_events, and leave the URLs out.
Sources
Related posts
More in Developers
- GET /v1/jobs thread_id filter: why a teammate's job still 404s
The thread_id filter on the Sume jobs list narrows results and never widens what an API key can read. Teammates' jobs stay 404, and only the creator can cancel.
- jobs_result batch: read a wave when some jobs are still running
Sume's batch jobs_result returns one entry per id in request order. Read ok per entry, treat job_not_completed as running, and re-read only failed_job_ids.
- jobs_wait wait_for any: other Sume jobs keep running and billing
jobs_wait takes up to 20 job_ids and wait_for all or any. With any it returns on the first finish, but the other jobs keep running and billing.
- A kill switch for paid Sume submits: stop new jobs, cancel queued
Add an off switch to code that spends on the Sume API: check a flag before each submit, then cancel queued jobs; a 409 job_generation_already_started will bill.
Written by Sume