Which Sume job and run endings send a webhook, and which stay silent?

Jobs send job.completed, job.failed and job.canceled. Canceled or skipped runs send nothing. A matrix of terminal states and what your receiver can expect.

5 min readSume
All posts

A generation job sends a webhook for all three endings: job.completed, job.failed and job.canceled. A Format, Action or Agent Completion run sends one webhook, *.run.terminal, when it completes or fails, and a canceled or skipped run sends nothing. If your receiver only fires on a callback, a canceled run looks like a run that never ended.

The matrix

Sume sends terminal events only. There are no progress or partial deliveries on either family, so a webhook is always the end of the story and never a status update. That also means a missing webhook can mean three different things: the work is still running, the ending is one that stays silent, or the delivery failed.

Poll status as the tiebreaker, because a status read answers all three questions with one request.

Terminal states and their webhooks (read 2026-10-07)
SurfaceEndingWebhook eventYour fallback
Jobcompletedjob.completedGET /v1/jobs/{id}/result
Jobfailedjob.failed with an error objectGET /v1/jobs/{id}
Jobcanceledjob.canceled with an error objectPoll status
Format runcompleted or failedformat.run.terminalGET /v1/format-runs/{id}
Format runcanceled or skippedNonePoll status or your cancel call's own response
Action runterminalaction.run.terminalGET /v1/action-runs/{id}
Agent Completioncompleted or failedagent.run.terminalGET /v1/agent-runs/{id}

What a run receipt adds

A run receipt carries an outcome of ok, degraded or error, and a run that finishes with partial output can arrive as degraded instead of failing. Treat degraded as a result to inspect, not as success and not as failure. Receipts bigger than 1 MiB arrive with payload: null and error.code set to payload_too_large. The run is fine in that case, and you fetch the full result from result_url.

Job webhooks carry the same ids you already have: request_id and job_id are the same value, and the docs tell you to use job_id as your idempotency key. Run receipts use request_id for that purpose, and it stays the same on every retry of the same run.

The event name is the cleanest router. Look at event first, then at status and outcome, then at the payload. A receiver that first looks for artifacts and then guesses the ending will mishandle job.failed and job.canceled, whose bodies carry an error object where the completed body carries artifacts.

Handle the silence yourself

Because canceled and skipped runs send nothing, write down who canceled. If your own code called the cancel endpoint, mark the row canceled in the same transaction, so there is nothing to wait for, and never rely on a callback to learn about your own action. If an operator or a rule can cancel from the dashboard, add a reconciler that reads runs which have been pending longer than you expect and polls them once.

A Format run has a built-in bound for that reconciler: its expires_at is 90 minutes after created_at, or earlier if the run goes silent, and after that Sume force-finalizes it as failed. A pending row older than that is something to investigate, not something to keep waiting on.

A delivery is not an outcome

The delivery has its own status, with pending, delivering, delivered, retrying, failed and exhausted. After ten refused attempts the delivery is exhausted, but the run is still completed, and you can replay the real terminal event: POST /v1/jobs/{id}/webhook/redeliver needs jobs:write, and POST /v1/format-runs/{id}/webhook/redeliver needs formats:write. A redeliver signs with the same secret and a fresh timestamp, so your verifier needs no change.

Design the receiver so that either path leads to the same state: a signed webhook and a status poll both write the same terminal row, once. A unique key on the job or run id, with a terminal status that only moves forward, gives you that without locks.

Return a 2xx for every signed event you recorded, including the endings you did not expect. A non-2xx response is a failed attempt, a 3xx is not followed, and each attempt has ten seconds, so a receiver that throws on job.canceled will have its own retries to deal with for nothing.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume