Mastra 1.72 crash recovery and leases: checkpoint the Sume job id

Mastra 1.72 adds multi-worker task leases and crash recovery. Store the Sume job id before waiting so a recovered worker polls instead of paying twice.

4 min readSume
All posts

Write the Sume job id to durable storage before any worker waits on it, then make recovery a status read instead of a second submit. The Mastra releases page lists 1.72.0 with multi-worker background-task ownership leases and crash recovery, which decide who resumes a task, but only your checkpoint decides whether the resumed task spends again.

Sume jobs are durable on its side: a valid submit returns a job id that survives your process restarting, so the only way to pay twice is to submit twice.

What the release page lists

Relevant items only.

Mastra releases (read 2026-10-03)
ReleaseItem
1.72.0Live channel resolvers
1.72.0Multi-worker background-task ownership leases
1.72.0Crash recovery
1.72.0Teams tools

The failure window

A worker can die in three places: before submit, after submit but before it saves the job id, and after it saved the id. Only the middle one is dangerous, and it is closed by the idempotency key rather than by the lease.

Where a crash can land
Crash pointStateRecovery
Before submitNo job existsSubmit with the stored key
After submit, before saveA job exists, id unknownResubmit with the same Idempotency-Key; Sume returns the original job
After saveJob id storedRead GET /v1/jobs/{id}/status; never submit

Derive the key from the task, not the attempt

Build the Idempotency-Key from something stable in your own data, such as an order, scene and revision, and not from a random value made when the step runs. A recovered worker then recomputes the same key. Sume's docs say to reuse a key only for the same operation and payload; a different payload under the same key returns 409 idempotency_conflict.

const key = `order-${orderId}-scene-${sceneIndex}-v${revision}`;

const res = await fetch("https://api.sume.com/v1/image-1.0/generate", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SUME_API_KEY}`,
    "Content-Type": "application/json",
    "Idempotency-Key": key,
  },
  body: JSON.stringify({ prompt, mode: "async" }),
});
const submitted = await res.json();
// save submitted.request_id with the task before awaiting anything

What the recovered worker does

On resume, branch on what the checkpoint holds.

  • Job id present: poll status_url with backoff until terminal is true, then fetch the result when result_ready is true.
  • Job id absent: recompute the key and submit again; the replay returns the original job.
  • Status failed or canceled: read the public error from the job record and decide on a new key, since a new key is a new paid job.

Test the recovery path

Kill the worker at each of the three points in the table above and check the outcome: one job, one charge. Sume's usage ledger records reserved, captured and refunded rows per job, so a duplicate would show up as a second reservation. Look up the cost of one task with GET /v1/usage?job_id=..., which also accepts run_id and thread_id scopes.

Run the same test with the lease handed to a second worker mid-wait; both should arrive at the same job id.

Leases and cancellation

If two workers briefly both believe they own a task, the idempotency key keeps them from creating two jobs. Cancellation is a separate decision: it succeeds only before generation starts, and afterwards returns 409 job_generation_already_started with the job running to completion.

A lease expiring is not a reason to cancel. Treat the stored job id as the source of truth for what was paid for.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume