Hatchet durable tasks: checkpoint a Sume job id and never pay twice

A Hatchet durable task replays from its last checkpoint after a crash. Make the Sume submit step safe with one Idempotency-Key, then wait on the job id.

5 min readSume
All posts

The short answer

Hatchet's durable execution page says each completed piece of a durable task creates a checkpoint in a durable event log, and after a worker crash the task replays from the last checkpoint without re-running earlier steps. Put the Sume submit in its own step, return the job id from it, and the paid call is not repeated after a crash that happens later.

The gap is a crash between the HTTP request leaving and the checkpoint being written. For that one window, send an Idempotency-Key that you derive from your business intent, so a repeated submit returns the original Sume job instead of creating a second one.

What Hatchet promises and what it leaves to you

The same page describes durable tasks as doing one of two things: waiting for something such as a sleep or an event, or spawning child tasks. It also names tasks that are hard to make idempotent, like sending an email, as a good fit because completed work is not replayed. A paid generation call is the same kind of side effect.

Checkpoints protect steps that finished. They cannot tell you whether a request that was in flight at the crash reached the provider. That is the part the Sume header covers.

Where each guard applies (read 2026-10-03)
Failure pointHatchet checkpointSume Idempotency-Key
Crash after the submit step is recordedReplay skips the stepNot needed
Crash after the request left, before it was recordedStep runs againRepeat returns the original job
Crash while waiting for the jobReplay resumes from the saved job idJob keeps running on Sume
Payload changed between attemptsNot detectedTreat as a new operation, use a new key

Shape of the workflow

Three steps keep the boundaries clean. Step one builds the key from the order and revision, submits with mode: async, and returns data.job.id plus data.idempotency_hit. Step two waits, either on a webhook-driven event or by sleeping for the interval in next_poll_after_seconds and reading the job status. Step three fetches the result only when result_ready is true; before that, /result answers 409 job_not_completed.

Keep the job id as the only state you pass between steps. It is small, it is stable, and anyone can re-read the job from it, which is what makes a replay cheap.

The submit step

This is the body of that first step as plain Python with the standard library. It runs on its own; wrap it in your Hatchet task function.

import asyncio, json, os, urllib.request

def submit(order_id: str, revision: int) -> str:
    key = f"{order_id}-hero-r{revision}"
    req = urllib.request.Request(
        "https://api.sume.com/v1/images",
        data=json.dumps({"model": "sume/auto", "mode": "async",
                         "prompt": "A ceramic mug on a pale oak table"}).encode(),
        headers={"authorization": f"Bearer {os.environ['SUME_API_KEY']}",
                 "idempotency-key": key,
                 "content-type": "application/json"},
        method="POST",
    )
    with urllib.request.urlopen(req) as r:
        return json.load(r)["data"]["job"]["id"]

async def main():
    print(await asyncio.to_thread(submit, "order-8823", 1))

asyncio.run(main())

Do not mint the key inside the retry

If the key comes from a clock or a random value generated inside the step, a replay produces a new one and the protection is gone. Derive it from inputs the workflow already holds, and bump the revision only when the prompt really changes. Sume's error guidance says to back off on 429 and use retry-after when present, so let the engine's retry policy honour that rather than looping inside the step.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume