BullMQ retry: exponential backoff for paid API jobs
Set attempts and an exponential backoff on BullMQ jobs, stop early with UnrecoverableError, and key each paid API call to the job so retries replay.

To retry a BullMQ job, add it with attempts greater than 1 and a backoff such as { type: 'exponential', delay: 1000 }: when the processor throws, BullMQ retries it after 2^(attempts − 1) × delay milliseconds, optionally with jitter, and without a backoff it retries at once. For a job that calls a paid API, build the Idempotency-Key from the job, so every attempt replays the same run, and throw UnrecoverableError on errors a retry can't fix.
BullMQ facts come from its Retrying failing jobs guide and the other BullMQ pages under Sources; Sume facts come from Create a run and Errors and spend. All were read on 2026-09-28. Sume has no BullMQ package: the worker makes one plain HTTPS call. For retrying a single request inside one process instead of the whole job, see Axios retry: retry a POST safely with an idempotency key.
How do I set attempts and backoff?
Pass them when you add the job, or once in the queue's defaultJobOptions. A job fails when its processor throws, and attempts counts the first try: attempts: 3 means at most 2 retries.
| Setting | What BullMQ does |
|---|---|
No backoff | Retries as soon as the job fails, with no delay |
{ type: 'fixed', delay } | Waits delay milliseconds before every retry |
{ type: 'exponential', delay } | Waits 2^(attempts − 1) × delay: with 1000, retries 1, 2 and 4 seconds apart |
jitter | 0 to 1, default 0. With 0.5, each wait is random between half and all of the computed delay |
backoffStrategy in worker settings | Custom delay per attempt; returning -1 moves the job to failed without a retry |
throw new UnrecoverableError() | Moves the job to failed with no retries, overriding attempts |
import { Queue } from "bullmq";
const queue = new Queue("videos", { connection });
await queue.add(
"promo",
{ orderId: "1042" },
{
jobId: "order-1042-promo", // a second add is ignored while this id is in the queue
attempts: 5,
backoff: { type: "exponential", delay: 1000, jitter: 0.5 },
},
);How do I make a retried job safe to pay for?
Give the job a custom jobId from your own record. BullMQ ignores a second add with an id that is still in the queue, and the id follows the job through every state, so a key built from it is the same on every attempt, and Sume answers a repeat with 200, the original receipt and idempotency_hit: true instead of a second charge. Then let Sume's error envelope decide: retryable says whether resending the same request can succeed. The docs say to retry a 503 later with the same key, and a 429 carries retry-after in seconds, which BullMQ's manual rate limit, queue.rateLimit() in milliseconds, can honor for the whole queue.
import { Worker, RateLimitError, UnrecoverableError } from "bullmq";
new Worker("videos", async (job) => {
const res = await fetch("https://api.sume.com/v1/formats/acme/product-promo/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": `${job.id}-v1`, // same on every attempt
},
body: JSON.stringify({
input: { order_id: job.data.orderId },
communication: { webhook_url: "https://example.com/hooks/sume" },
}),
});
const body = await res.json();
if (res.ok) return body.data.id; // 202: new run, 200: replay
if (res.status === 429) {
await queue.rateLimit(Number(res.headers.get("retry-after")) * 1000); // ms
throw new RateLimitError(); // back to waiting, not a failed attempt
}
if (body.error.retryable || res.status === 503) throw new Error(body.error.code);
throw new UnrecoverableError(`${res.status} ${body.error.code}`);
}, { connection, limiter: { max: 1, duration: 500 } });Which errors should stop retrying?
Most 4xx answers. A 4xx at create means nothing ran and nothing was charged, so the worker throws UnrecoverableError and the job fails at once instead of spending its attempts. Each Sume error carries a next_action (fix_input, authenticate, add_funds, retry_later, …) that says which case it is in; Axios retry: retry a POST safely lists the codes a retry can't fix, from 402 insufficient_credits to a 502 that is really your input.
The answers Sume marks retryable, such as 409 idempotency_key_in_use while the first attempt is still in flight, and a 503 become plain thrown errors, so they get the job's backoff. A 429 is different: a job rate limited with RateLimitError doesn't use up attempts, because BullMQ doesn't count rate limiting as a real error. To cap it anyway, BullMQ's docs check job.attemptsStarted against job.opts.attempts and throw UnrecoverableError.
What if the run fails after the job succeeded?
Then a BullMQ retry can't help: the old key is bound to the failed receipt, so the same key only replays the failure. Start a new job with a new jobId, such as order-1042-promo-2, so its key is new too. Don't make the processor wait for the video either: long-form video is 15 to 30 minutes of work, and a worker that stops waiting doesn't stop the run or its spend. Return the run id and let Sume POST one signed format.run.terminal receipt to communication.webhook_url when the run completes or fails.
Sources
Related posts
More in Integrations
- Claude Code: allow MCP tools without approving every call
Allow MCP tools in Claude Code with permission rules named mcp__server__tool. Allow one tool or a whole server; keep paid tools on ask.
- Claude Code MCP project scope: share a server with your team
Claude Code's project scope saves an MCP server to .mcp.json at the repo root for the team to commit. How approval, sign-in, and keys work.
- Cloud Scheduler trigger for a Cloud Run job, retry-safe
Add a Cloud Scheduler trigger to a Cloud Run job with a cron and a time zone. A failed task retries 3 times by default, so key paid calls to the date.
- Codex MCP tool timeout: tool_timeout_sec and slow jobs
Codex gives each MCP tool call 60 seconds by default. Raise tool_timeout_sec per server in config.toml, or keep slow media jobs inside the limit.
Written by Sume