Celery task retry: backoff and jitter for a paid API call
Retry a Celery task with autoretry_for, retry_backoff and max_retries, wait out retry-after on a 429, and send one Idempotency-Key on every retry.

To retry a Celery task, list the exceptions to retry in autoretry_for, or catch the error and raise self.retry(exc=exc, countdown=60); add retry_backoff=True to delay those automatic retries 1, 2, 4, 8 seconds…, with jitter randomizing each delay, and set max_retries, which defaults to 3. When the task calls a paid API, retry only the errors a retry can fix, and send the same Idempotency-Key on every attempt, so a retry returns the original run instead of paying for a second one.
Celery facts come from its Tasks guide, Requests facts from its Quickstart, and Sume facts from Create a run and Errors and spend. All were read on 2026-09-28. Sume has no Celery package: the task makes one plain HTTPS call. Whether a video job needs a Celery worker at all is covered in FastAPI long running task: return 202 with a job id.
What retry options does a Celery task have?
Set them on the task decorator. retry_backoff, retry_backoff_max and retry_jitter shape automatic retries from autoretry_for; a manual self.retry() waits default_retry_delay unless you pass countdown. Either way, Celery sends a new message with the same task id.
| Option | Default | What it does |
|---|---|---|
autoretry_for | No exceptions | Retries the task when one of these exception classes is raised |
max_retries | 3 | Maximum retries before giving up; None retries forever |
retry_backoff | False: autoretries are not delayed | True waits 1, 2, 4, 8 seconds…; a number is the delay factor |
retry_backoff_max | 600 seconds | Caps the backoff delay |
retry_jitter | True | Picks a random delay between zero and the backoff value |
default_retry_delay | 3 minutes | Delay for self.retry(); countdown overrides it per call |
How do I retry a paid API call safely?
Build the key from the task's arguments, the thing being made, never from uuid4() per attempt: Sume's docs say a per-request key makes the header decorative. Then let Sume's error envelope decide. Its retryable flag says whether resending the same request can succeed, a 429 carries a retry-after header with the seconds to wait, and the docs say to retry a 503 later with the same key. Requests raises ConnectionError on a network problem and Timeout when the server stops answering within timeout.
import os, requests
from celery import shared_task
class Transient(Exception): pass
@shared_task(bind=True, max_retries=5, retry_backoff=True,
autoretry_for=(requests.ConnectionError, requests.Timeout, Transient))
def start_video(self, order_id):
r = requests.post(
"https://api.sume.com/v1/formats/acme/product-promo/runs",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": f"order-{order_id}-promo-v1", # same on every retry
},
json={"input": {"order_id": order_id},
"communication": {"webhook_url": "https://example.com/hooks/sume"}},
timeout=30,
)
if r.ok:
return r.json()["data"]["id"] # 202: new run, 200: replay
err = r.json()["error"]
if r.status_code == 429:
raise self.retry(countdown=int(r.headers["retry-after"]))
if err["retryable"] or r.status_code == 503:
raise Transient(err["code"]) # autoretried with backoff and jitter
raise RuntimeError(f"{r.status_code} {err['code']}") # fix the call insteadWhich errors should the task retry, and which should fail?
Retry timeouts, dropped connections, and the answers Sume marks retryable, such as 409 idempotency_key_in_use while the first attempt is still in flight, plus a 429 after retry-after and a 503 studio_agent_upstream_unavailable later. A timeout doesn't tell you whether the create arrived, and the key makes the resend safe: Sume answers the same key and body with 200, the original receipt and idempotency_hit: true, so nothing is charged twice.
Fail everything else at once. A 4xx at create means nothing ran and nothing was charged, so the code raises a RuntimeError, which isn't in autoretry_for, and Celery doesn't retry it. Axios retry: retry a POST safely lists the codes a retry can't fix, from 402 insufficient_credits to a 502 that is really your input. When max_retries is used up, the task fails too: Celery re-raises the current exception, or raises MaxRetriesExceededError when self.retry() got no exc, as in the 429 branch.
What if the run itself fails or takes 20 minutes?
A run that ended failed is a different case: the old key is bound to that receipt, so a same-key retry only replays the failure. Retry it with a new key, such as order-1042-promo-v2. And don't keep a worker waiting for the video: long-form video is 15 to 30 minutes of work, and a worker that stops waiting doesn't stop the run or its spend. Return the run id, store it, and let Sume POST one signed format.run.terminal receipt to communication.webhook_url when the run completes or fails.
With the key in place the task is idempotent, the condition Celery's docs set before you enable acks_late, which acknowledges the message after the task returns instead of before it starts.
Sources
Related posts
More in Integrations
- Claude Code: allow MCP tools without approving every call
Allow MCP tools in Claude Code with permission rules named mcp__server__tool. Allow one tool or a whole server; keep paid tools on ask.
- Claude Code MCP project scope: share a server with your team
Claude Code's project scope saves an MCP server to .mcp.json at the repo root for the team to commit. How approval, sign-in, and keys work.
- Cloud Scheduler trigger for a Cloud Run job, retry-safe
Add a Cloud Scheduler trigger to a Cloud Run job with a cron and a time zone. A failed task retries 3 times by default, so key paid calls to the date.
- Codex MCP tool timeout: tool_timeout_sec and slow jobs
Codex gives each MCP tool call 60 seconds by default. Raise tool_timeout_sec per server in config.toml, or keep slow media jobs inside the limit.
Written by Sume