What is a dead letter queue? DLQs for AI job pipelines

A dead-letter queue holds messages that failed processing too many times, so they stop looping and can be inspected. Which AI job failures go there.

5 min readSume
All posts

A dead-letter queue (DLQ) is a separate queue that receives messages your consumers failed to process after a set number of attempts. Moving them aside stops a bad message from being retried forever or blocking the main queue, and keeps it where someone can inspect it, fix the cause, and send it back.

The queue settings below come from Amazon's SQS dead-letter queue guide, read on 2026-09-28; other brokers name the same idea differently. The AI job examples use Sume's Errors and rate limits and Webhooks docs.

How does a dead-letter queue work?

In Amazon SQS, the source queue gets a redrive policy that names the DLQ and a maxReceiveCount: the number of times a consumer can receive a message before SQS moves it to the DLQ. Each failed attempt is a receive that didn't end in a delete. Once the count is reached, the message leaves the main queue.

Amazon lists what the DLQ is for once messages land there:

  • Examine logs for the exceptions that moved them.
  • Analyze the message contents to diagnose the application.
  • Check whether the consumer had enough time to process messages.
  • Move messages back out with dead-letter queue redrive.

How do I configure a dead-letter queue in SQS?

Create the DLQ first as an ordinary queue, then point the source queue's redrive policy at it. The settings that matter:

From Amazon's Using dead-letter queues in Amazon SQS, read 2026-09-28.
SettingWhat Amazon's guide says
maxReceiveCountReceives before a message moves; a value of 1 moves it after one failure, so set it high enough for retries
LocationThe DLQ must be in the same AWS account and Region as the source queue
Redrive allow policyAllows all source queues by default; byQueue names up to 10; denyAll blocks use as a DLQ
RetentionStandard queues keep the original enqueue time, so give the DLQ a longer retention period than the source
FIFO queuesDon't use a DLQ if you can't break the exact order of messages

Which failed AI jobs belong in a dead-letter queue?

The ones a retry won't fix. When a queued message submits a paid generation job, the job's error category says which kind of failure it was. Poll the job until its status is terminal before you judge it; Sume's job statuses end at completed, failed or canceled (Errors and rate limits).

  • Dead-letter them: validation (fix input), auth (check the API key and workspace access), quota (add funds or lower request cost), and generation_rejected (inspect events and fix unsupported input). Retrying the same message repeats the same failure.
  • Retry them later from the main queue: queue, generation_unavailable, and runtime_unavailable (retry later; do not retry aggressively). A replay with the same Idempotency-Key returns the original job (Video generation), so resubmitting a job that already failed takes a new key.
  • Check before retrying: for generation_timeout and worker_timeout the docs say to poll status or retry later, so read the job's status before you resubmit.
  • Keep internal failures with their request and job ids: the docs' next action is to contact support.

Is a failed webhook delivery a dead letter?

No. On Sume, a job webhook is retried up to 10 attempts in total, and ten refused attempts leave a failed delivery and a job that still reached its real terminal state (Webhooks). The job didn't fail, so don't resubmit it. POST /v1/jobs/{job_id}/webhook/redeliver re-sends the terminal event even after automatic attempts are exhausted, and your receiver should treat job_id as the idempotency key. Sume webhook not received? walks through the checks.

Do failed jobs in the DLQ cost money?

Sume's docs say provider-backed generation can reserve estimated usage, capture actual cost on success, and refund on failure or cancellation before capture (Core workflow). Do failed AI video jobs cost money? covers the details. What a DLQ saves you is the cost of resubmitting a job that will fail again.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume