What is a dead letter queue? DLQs for AI job pipelines
A dead-letter queue holds messages that failed processing too many times, so they stop looping and can be inspected. Which AI job failures go there.

A dead-letter queue (DLQ) is a separate queue that receives messages your consumers failed to process after a set number of attempts. Moving them aside stops a bad message from being retried forever or blocking the main queue, and keeps it where someone can inspect it, fix the cause, and send it back.
The queue settings below come from Amazon's SQS dead-letter queue guide, read on 2026-09-28; other brokers name the same idea differently. The AI job examples use Sume's Errors and rate limits and Webhooks docs.
How does a dead-letter queue work?
In Amazon SQS, the source queue gets a redrive policy that names the DLQ and a maxReceiveCount: the number of times a consumer can receive a message before SQS moves it to the DLQ. Each failed attempt is a receive that didn't end in a delete. Once the count is reached, the message leaves the main queue.
Amazon lists what the DLQ is for once messages land there:
- Examine logs for the exceptions that moved them.
- Analyze the message contents to diagnose the application.
- Check whether the consumer had enough time to process messages.
- Move messages back out with dead-letter queue redrive.
How do I configure a dead-letter queue in SQS?
Create the DLQ first as an ordinary queue, then point the source queue's redrive policy at it. The settings that matter:
| Setting | What Amazon's guide says |
|---|---|
maxReceiveCount | Receives before a message moves; a value of 1 moves it after one failure, so set it high enough for retries |
| Location | The DLQ must be in the same AWS account and Region as the source queue |
| Redrive allow policy | Allows all source queues by default; byQueue names up to 10; denyAll blocks use as a DLQ |
| Retention | Standard queues keep the original enqueue time, so give the DLQ a longer retention period than the source |
| FIFO queues | Don't use a DLQ if you can't break the exact order of messages |
Which failed AI jobs belong in a dead-letter queue?
The ones a retry won't fix. When a queued message submits a paid generation job, the job's error category says which kind of failure it was. Poll the job until its status is terminal before you judge it; Sume's job statuses end at completed, failed or canceled (Errors and rate limits).
- Dead-letter them:
validation(fix input),auth(check the API key and workspace access),quota(add funds or lower request cost), andgeneration_rejected(inspect events and fix unsupported input). Retrying the same message repeats the same failure. - Retry them later from the main queue:
queue,generation_unavailable, andruntime_unavailable(retry later; do not retry aggressively). A replay with the sameIdempotency-Keyreturns the original job (Video generation), so resubmitting a job that already failed takes a new key. - Check before retrying: for
generation_timeoutandworker_timeoutthe docs say to poll status or retry later, so read the job's status before you resubmit. - Keep
internalfailures with their request and job ids: the docs' next action is to contact support.
Is a failed webhook delivery a dead letter?
No. On Sume, a job webhook is retried up to 10 attempts in total, and ten refused attempts leave a failed delivery and a job that still reached its real terminal state (Webhooks). The job didn't fail, so don't resubmit it. POST /v1/jobs/{job_id}/webhook/redeliver re-sends the terminal event even after automatic attempts are exhausted, and your receiver should treat job_id as the idempotency key. Sume webhook not received? walks through the checks.
Do failed jobs in the DLQ cost money?
Sume's docs say provider-backed generation can reserve estimated usage, capture actual cost on success, and refund on failure or cancellation before capture (Core workflow). Do failed AI video jobs cost money? covers the details. What a DLQ saves you is the cost of resubmitting a job that will fail again.
Sources
Related posts
More in Developers
- What is a video API? The five kinds, explained
A video API lets code make, edit or deliver video over HTTP. The five kinds, what each one takes and returns, and how to tell which one you need.
- Edit decision list (EDL): what it is, with an example
An edit decision list (EDL) is the ordered list of edits that rebuilds a cut: source, track, transition, and timecodes. An example and its JSON form.
- What is serverless inference? How it works and is billed
Serverless inference means calling a hosted AI model over HTTP with no servers of your own. Providers bill for compute time or for each output.
- yuv420p10le vs yuv420p: 8-bit vs 10-bit pixel formats
yuv420p and yuv420p10le are both planar YUV 4:2:0. yuv420p stores 8 bits per sample; yuv420p10le stores 10, little-endian, in 16-bit words.
Written by Sume