n8n durable agent message queue: replays and Sume idempotency keys

n8n 2.42 lays a durable agent message queue foundation. A queue that can redeliver means a paid Sume call needs a stable idempotency_key per message.

5 min readSume
All posts

A durable queue exists so a message survives a crash and is handled again, which means any paid Sume call made while handling it must be safe to run twice. Make the idempotency_key a function of the message, not of the attempt. n8n 2.42.0 (2026-09-29) lists a durable agent message queue foundation among its notable changes (read 2026-10-03); the release notes give no detail beyond that line, so this post covers the Sume side of the contract rather than n8n's internals.

The n8n item is in the 2.42.0 release notes. The Sume gates are in MCP tools and gates and the job behavior in Jobs and results.

Why message-level keys

At-least-once delivery is the usual deal with durable queues: a message is retried until the consumer acknowledges it. If the consumer crashed after Sume accepted a paid create but before the acknowledgment, the retry sends the create again. With a key derived from the message id, Sume sees the same idempotency_key and returns the original job. With a key derived from a timestamp or a random value, Sume sees a new request and starts a second job.

Sume documents the key as a stable key for transport and dedup, not human approval. For the run endpoints, a replayed Idempotency-Key returns the original receipt with idempotency_hit: true, and reusing a key with a different payload returns 409 idempotency_conflict.

Key sources compared, read 2026-10-03
Key built fromOn a retryResult
Queue message idSameOriginal job returned
Message id plus item indexSame per itemEach item once
TimestampDifferentA second paid job
Random value per attemptDifferentA second paid job

Build the key from the message

A message may produce several paid calls, so add the item index. The helper below is deterministic: the same message and index always give the same key, and anything else gives a different one. It also refuses an empty message id, which would collapse unrelated messages into one key.

import hashlib

def key_for(message_id: str, item: int, purpose: str = "gen") -> str:
    if not message_id:
        raise ValueError("message_id is empty")
    h = hashlib.sha256(f"{message_id}|{item}".encode()).hexdigest()[:16]
    return f"{purpose}-{h}"

if __name__ == "__main__":
    print(key_for("msg_42", 0), key_for("msg_42", 0), key_for("msg_42", 1))

Pair it with a cap and a wait

A key prevents duplicates, not overspending, so keep max_spend_usd on each call. After the create, wait with jobs_wait, which holds up to 55 seconds; if it returns wait_slice_expired, retry the wait, and never resubmit the paid create. For a message that fans out to many jobs, one batch wait on up to 20 ids is cheaper than one wait each.

Finally, make the acknowledgment the last step, after the job id is stored. If you acknowledge before storing it, a crash loses the id; if you store first and crash before acknowledging, the retry hits the same key and finds the job. That ordering turns a durable queue into an exactly-once effect for paid work.

Edge cases that break the key

Three edge cases are worth a test. First, a message whose payload is edited between attempts: Sume answers 409 idempotency_conflict on the run endpoints when a key is reused with a different payload, so an edited retry fails loudly instead of silently running something new. Decide whether your queue should then mint a new message id, which makes it a new job on purpose, or stop and ask.

Second, a very old message replayed after days. A key you reuse long after the original call may no longer behave as a replay, since Sume's docs do not promise unlimited key retention, so check job status first with jobs_list or jobs_status when a retry is old. Third, two consumers handling the same message at once, which a durable queue can allow after a timeout. The same key protects you there too, because both requests carry it, and Sume treats them as one.

Log the key with the job id so each of these cases is a lookup, not a puzzle.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume