Sidekiq retry: backoff, retry count and paid API calls

Sidekiq retries a failed job 25 times over about 20 days by default. Cap it, wait retry-after with sidekiq_retry_in, and reuse one Idempotency-Key.

5 min readSume
All posts

By default Sidekiq retries a failed job 25 times with exponential backoff, waiting (retry_count ** 4) + 15 seconds plus a random amount, over approximately 20 days, then moves it to the Dead set. For a job that starts paid work over an API, cap the count with sidekiq_options retry: 5, use sidekiq_retry_in to wait the API's retry-after or to kill errors a retry can't fix, and build the Idempotency-Key from the job's arguments, so a retry replays the original run instead of paying for a second one.

Sidekiq facts come from its Error Handling, Best Practices and The Basics wiki pages and Ruby's Net::HTTP docs; Sume facts come from Create a run, Errors and spend and Runs and results. All were read on 2026-09-28. Sume has no Sidekiq gem: the job makes one plain HTTPS call. The same pattern in Python is Celery task retry: backoff and jitter for a paid API call.

How long does Sidekiq keep retrying?

The full formula is (retry_count ** 4) + 15 + (rand(10) * (retry_count + 1)) seconds, so early retries come quickly and late ones days apart. After the last retry the job goes to the Dead set, which holds up to 10,000 jobs or 6 months; from there you retry it by hand in the Web UI, whose Retries and Dead tabs list failed jobs.

From Sidekiq's Error Handling wiki, which computes these times assuming rand(10) returns 5, read 2026-09-28.
RetryWait before itTotal wait so far
120s20s
54m 56s8m 24s
101h 50m 26s4h 22m 38s
1510h 41m 46s1d 11h 41m 52s
201d 12h 13m 56s6d 12h 40m 16s
253d 20h 11m 56s20d 10h 17m 0s

How do I change the retry count?

Per job class, with sidekiq_options:

  • retry: 5: five retries, then the Dead set.
  • retry: 0: no retries; a failed job goes straight to the Dead set.
  • retry: false: the job is discarded if it fails.
  • retry_for: 48.hours (Sidekiq 7.1.3 and later): retry for a period of time instead of a count.
  • :max_retries in sidekiq.yml sets the maximum globally.
  • sidekiq_retries_exhausted runs right before a job moves to the Dead set: a place to alert someone.

How do I retry a paid API call safely?

Send the same key and body on every run. Sidekiq's best practices say it executes a job at least once, not exactly once, and even a completed job can be re-run. The arguments travel inside the job's JSON hash, so a key built from them is the same every time, and Sume answers a repeat with 200, the original receipt and idempotency_hit: true instead of a second charge. A random key per attempt would, in the words of Sume's docs, make the header decorative; Idempotency keys for AI video APIs covers the other replay cases.

sidekiq_retry_in receives the retry count and the exception: return seconds to wait, :kill to send the job to the Dead set (Sidekiq 6.5.2 and later), or nil for the default delay. Sume's error envelope says which case applies: retryable tells whether resending can succeed, a 429 carries retry-after in seconds, and the docs say to retry a 503 later with the same key. Any other 4xx that isn't marked retryable means the call itself needs fixing: nothing ran and nothing was charged, which is why Fatal goes straight to the Dead set instead of through five retries. In Ruby's Net::HTTP, res.code is a string such as "429", and res["retry-after"] reads the header.

require "net/http"
class StartVideoJob
  include Sidekiq::Job
  sidekiq_options retry: 5
  RateLimited = Class.new(StandardError)
  Fatal = Class.new(StandardError)
  sidekiq_retry_in do |count, exception|
    case exception
    when RateLimited then exception.message.to_i # the API's retry-after
    when Fatal then :kill # to the Dead set, no more retries
    end # nil: Sidekiq's default backoff
  end
  def perform(order_id, version = 1)
    res = Net::HTTP.post(URI("https://api.sume.com/v1/formats/acme/product-promo/runs"),
      { input: { order_id: order_id },
        communication: { webhook_url: "https://example.com/hooks/sume" } }.to_json,
      { "Authorization" => "Bearer #{ENV.fetch("SUME_API_KEY")}", "Content-Type" => "application/json",
        "Idempotency-Key" => "order-#{order_id}-promo-v#{version}" }) # same on every retry
    body = JSON.parse(res.body)
    return body["data"]["id"] if res.is_a?(Net::HTTPSuccess) # 202 new, 200 replay: store it
    raise RateLimited, res["retry-after"] if res.code == "429"
    raise body["error"]["code"] if body["error"]["retryable"] || res.code == "503"
    raise Fatal, "#{res.code} #{body["error"]["code"]}"
  end
end

What if the video run fails after the job has finished?

A Sidekiq retry can't fix that: the old key is bound to the failed receipt, so reusing it only replays the failure. Enqueue a new job with the next version, StartVideoJob.perform_async(1042, 2), whose key ends in -v2. Don't hold a Sidekiq thread waiting for the video either; long-form video is 15 to 30 minutes of work, and a job that stops waiting doesn't stop the run or its spend. Sume POSTs one signed format.run.terminal receipt to communication.webhook_url when the run completes or fails; Rails webhook HMAC signature verification covers receiving it.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume