Mastra 1.72 crash recovery and leases: checkpoint the Sume job id
Mastra 1.72 adds multi-worker task leases and crash recovery. Store the Sume job id before waiting so a recovered worker polls instead of paying twice.

Write the Sume job id to durable storage before any worker waits on it, then make recovery a status read instead of a second submit. The Mastra releases page lists 1.72.0 with multi-worker background-task ownership leases and crash recovery, which decide who resumes a task, but only your checkpoint decides whether the resumed task spends again.
Sume jobs are durable on its side: a valid submit returns a job id that survives your process restarting, so the only way to pay twice is to submit twice.
What the release page lists
Relevant items only.
| Release | Item |
|---|---|
| 1.72.0 | Live channel resolvers |
| 1.72.0 | Multi-worker background-task ownership leases |
| 1.72.0 | Crash recovery |
| 1.72.0 | Teams tools |
The failure window
A worker can die in three places: before submit, after submit but before it saves the job id, and after it saved the id. Only the middle one is dangerous, and it is closed by the idempotency key rather than by the lease.
| Crash point | State | Recovery |
|---|---|---|
| Before submit | No job exists | Submit with the stored key |
| After submit, before save | A job exists, id unknown | Resubmit with the same Idempotency-Key; Sume returns the original job |
| After save | Job id stored | Read GET /v1/jobs/{id}/status; never submit |
Derive the key from the task, not the attempt
Build the Idempotency-Key from something stable in your own data, such as an order, scene and revision, and not from a random value made when the step runs. A recovered worker then recomputes the same key. Sume's docs say to reuse a key only for the same operation and payload; a different payload under the same key returns 409 idempotency_conflict.
const key = `order-${orderId}-scene-${sceneIndex}-v${revision}`;
const res = await fetch("https://api.sume.com/v1/image-1.0/generate", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify({ prompt, mode: "async" }),
});
const submitted = await res.json();
// save submitted.request_id with the task before awaiting anythingWhat the recovered worker does
On resume, branch on what the checkpoint holds.
- Job id present: poll
status_urlwith backoff untilterminalis true, then fetch the result whenresult_readyis true. - Job id absent: recompute the key and submit again; the replay returns the original job.
- Status
failedorcanceled: read the public error from the job record and decide on a new key, since a new key is a new paid job.
Test the recovery path
Kill the worker at each of the three points in the table above and check the outcome: one job, one charge. Sume's usage ledger records reserved, captured and refunded rows per job, so a duplicate would show up as a second reservation. Look up the cost of one task with GET /v1/usage?job_id=..., which also accepts run_id and thread_id scopes.
Run the same test with the lease handed to a second worker mid-wait; both should arrive at the same job id.
Leases and cancellation
If two workers briefly both believe they own a task, the idempotency key keeps them from creating two jobs. Cancellation is a separate decision: it succeeds only before generation starts, and afterwards returns 409 job_generation_already_started with the job running to completion.
A lease expiring is not a reason to cancel. Treat the stored job id as the source of truth for what was paid for.
Sources
Related posts
More in Integrations
- n8n 2.41.6 task runner and a Sume webhook verifier that never throws
n8n 2.41.6 keeps its task runner alive on unhandled rejections. Write the Sume webhook check so a bad signature returns false; answer 2xx only after storing.
- OpenAI Agents SDK 0.23 MCP listing limits and Sume tools
openai-agents-python 0.23 adds configurable MCP listing page limits. What that means for Sume's hosted MCP, where the tool list depends on your OAuth scope.
- A Supabase job table for Sume webhooks: upsert on job_id
Sume retries webhooks up to 10 times. A Postgres table keyed on job_id with ON CONFLICT turns repeat deliveries into no-ops. Schema, SQL and the order of steps.
- Vercel AI SDK tool search maxResults and Sume tool groups
ai@7.0.127 tool search ranks deferred tools with a search() callback and maxResults. How to split Sume's hosted MCP tools into always-on and deferred groups.
Written by Sume