SQS FIFO 5-minute dedup vs a Sume Idempotency-Key: which wins?

SQS FIFO dedupes for only 5 minutes. A Sume Idempotency-Key covers the create call itself, so use both when a worker can retry after the window.

4 min readSume
All posts

They protect different hops, so use both. SQS FIFO deduplicates messages sent within a 5 minute interval, using either a MessageDeduplicationId or, with content-based deduplication, a SHA-256 of the body. That stops duplicate enqueues. It does not stop a worker that crashes after calling Sume and gets the message again, and it does not help after the window. A Sume Idempotency-Key does that job at the API.

What each layer covers

Deduplication layers (read 2026-10-02)
LayerWhat it dedupesLimit
SQS FIFO deduplicationDuplicate sends to the queue5 minute interval
Consumer retry after crashNot covered by SQS dedupVisibility timeout redelivery
Sume Idempotency-KeyDuplicate create callsSame key and body returns the original

Where duplicates still appear

Suppose a Lambda or container reads a message, calls Sume, and times out before deleting the message. The message becomes visible again and a second worker repeats the call. SQS dedup is irrelevant here, because the message was sent once. With a fixed Idempotency-Key built from the message's business id, the second call returns the original job instead of creating another paid one.

Conversely, if a producer re-sends the same business event ten minutes later, SQS will accept it as new. The Sume key still matches, provided the body is the same.

Pick the key carefully

Use the same value for the SQS dedup id and the Sume key if you can: one identity per business event. Do not use a random UUID created in the consumer, since each retry would invent a new one.

  • Same key plus a different body returns 409 idempotency_conflict; treat it as a bug and alert.
  • On 429 rate_limited honour retry-after; on 402 insufficient_credits do not retry.
  • For Format runs the replay returns 200 with idempotency_hit: true, so check for it instead of assuming a 202.

Finish the loop with a webhook

After a 202 the consumer should delete the message and store the job id. Completion comes back through job.completed, job.failed or job.canceled, deduped on job_id. Do not keep the SQS message invisible while a video renders; the generation is not tied to your queue's timeout, and a client timeout does not cancel a running job.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume