Celery task retry: backoff and jitter for a paid API call

Retry a Celery task with autoretry_for, retry_backoff and max_retries, wait out retry-after on a 429, and send one Idempotency-Key on every retry.

5 min readSume
All posts

To retry a Celery task, list the exceptions to retry in autoretry_for, or catch the error and raise self.retry(exc=exc, countdown=60); add retry_backoff=True to delay those automatic retries 1, 2, 4, 8 seconds…, with jitter randomizing each delay, and set max_retries, which defaults to 3. When the task calls a paid API, retry only the errors a retry can fix, and send the same Idempotency-Key on every attempt, so a retry returns the original run instead of paying for a second one.

Celery facts come from its Tasks guide, Requests facts from its Quickstart, and Sume facts from Create a run and Errors and spend. All were read on 2026-09-28. Sume has no Celery package: the task makes one plain HTTPS call. Whether a video job needs a Celery worker at all is covered in FastAPI long running task: return 202 with a job id.

What retry options does a Celery task have?

Set them on the task decorator. retry_backoff, retry_backoff_max and retry_jitter shape automatic retries from autoretry_for; a manual self.retry() waits default_retry_delay unless you pass countdown. Either way, Celery sends a new message with the same task id.

From Celery's Tasks guide, read 2026-09-28.
OptionDefaultWhat it does
autoretry_forNo exceptionsRetries the task when one of these exception classes is raised
max_retries3Maximum retries before giving up; None retries forever
retry_backoffFalse: autoretries are not delayedTrue waits 1, 2, 4, 8 seconds…; a number is the delay factor
retry_backoff_max600 secondsCaps the backoff delay
retry_jitterTruePicks a random delay between zero and the backoff value
default_retry_delay3 minutesDelay for self.retry(); countdown overrides it per call

How do I retry a paid API call safely?

Build the key from the task's arguments, the thing being made, never from uuid4() per attempt: Sume's docs say a per-request key makes the header decorative. Then let Sume's error envelope decide. Its retryable flag says whether resending the same request can succeed, a 429 carries a retry-after header with the seconds to wait, and the docs say to retry a 503 later with the same key. Requests raises ConnectionError on a network problem and Timeout when the server stops answering within timeout.

import os, requests
from celery import shared_task

class Transient(Exception): pass
@shared_task(bind=True, max_retries=5, retry_backoff=True,
             autoretry_for=(requests.ConnectionError, requests.Timeout, Transient))
def start_video(self, order_id):
    r = requests.post(
        "https://api.sume.com/v1/formats/acme/product-promo/runs",
        headers={
            "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
            "Idempotency-Key": f"order-{order_id}-promo-v1",  # same on every retry
        },
        json={"input": {"order_id": order_id},
              "communication": {"webhook_url": "https://example.com/hooks/sume"}},
        timeout=30,
    )
    if r.ok:
        return r.json()["data"]["id"]  # 202: new run, 200: replay
    err = r.json()["error"]
    if r.status_code == 429:
        raise self.retry(countdown=int(r.headers["retry-after"]))
    if err["retryable"] or r.status_code == 503:
        raise Transient(err["code"])  # autoretried with backoff and jitter
    raise RuntimeError(f"{r.status_code} {err['code']}")  # fix the call instead

Which errors should the task retry, and which should fail?

Retry timeouts, dropped connections, and the answers Sume marks retryable, such as 409 idempotency_key_in_use while the first attempt is still in flight, plus a 429 after retry-after and a 503 studio_agent_upstream_unavailable later. A timeout doesn't tell you whether the create arrived, and the key makes the resend safe: Sume answers the same key and body with 200, the original receipt and idempotency_hit: true, so nothing is charged twice.

Fail everything else at once. A 4xx at create means nothing ran and nothing was charged, so the code raises a RuntimeError, which isn't in autoretry_for, and Celery doesn't retry it. Axios retry: retry a POST safely lists the codes a retry can't fix, from 402 insufficient_credits to a 502 that is really your input. When max_retries is used up, the task fails too: Celery re-raises the current exception, or raises MaxRetriesExceededError when self.retry() got no exc, as in the 429 branch.

What if the run itself fails or takes 20 minutes?

A run that ended failed is a different case: the old key is bound to that receipt, so a same-key retry only replays the failure. Retry it with a new key, such as order-1042-promo-v2. And don't keep a worker waiting for the video: long-form video is 15 to 30 minutes of work, and a worker that stops waiting doesn't stop the run or its spend. Return the run id, store it, and let Sume POST one signed format.run.terminal receipt to communication.webhook_url when the run completes or fails.

With the key in place the task is idempotent, the condition Celery's docs set before you enable acks_late, which acknowledges the message after the task returns instead of before it starts.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume