Stripe retries webhooks for 3 days, Sume job webhooks for 4.5 minutes

Compare retry windows: Stripe live mode retries up to three days; Sume job webhooks use 10 attempts at 30s spacing. Size your poll fallback to the shorter one.

5 min readSume
All posts

If your webhook endpoint is down during a deploy, a Stripe event still arrives later: Stripe's webhook guide says it retries live-mode deliveries for up to three days with exponential backoff. A Sume job webhook does not give you that runway. Job webhooks get up to 10 attempts with a fixed delay of 30 seconds by default, so the automatic window is about 4.5 minutes. Run webhooks (Formats, Actions, Agent Completions) back off and last about three hours. After that the job or run is still finished, and you recover with a poll or a redeliver.

The arithmetic

Ten attempts have nine gaps between them. Job webhooks: 9 gaps x 30s = 270 seconds, or 4.5 minutes. Run webhooks use min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h). Without jitter the nine gaps are 30, 60, 120, 240, 480, 960, 1,920, 3,600 and 3,600 seconds, which sum to 11,010 seconds, or 3 hours 3 minutes 30 seconds. Each attempt also has a 10-second timeout, and Retry-After on a 429 or 503 can stretch a gap, so treat both totals as approximate.

Retry windows, from the vendor pages (read 2026-10-05)
DeliveryAttemptsSpacingApproximate window
Stripe live modeUntil about three daysExponential backoffUp to 3 days
Sume job webhook10Fixed, 30s by default270 seconds
Sume run webhook1030s x 2^(n-1), capped at 1hAbout 3h 3m 30s before jitter

What to do about the short window

A short automatic window is fine when recovery is cheap. Sume's docs say delivery is an optimization and the status_url polls stay in place. Manual redeliver exists too: POST /v1/jobs/{job_id}/webhook/redeliver (scope jobs:write) re-sends the real terminal event with a fresh timestamp and signature, and it still works after the automatic attempts are used up.

  • Run a sweeper every few minutes over jobs you submitted but never saw a terminal event for.
  • Poll with exponential backoff and stop on completed, failed or canceled.
  • Redeliver after a deploy freeze instead of resubmitting. A resubmit of the paid request is a second charge.
  • Alert on webhook_delivery status exhausted, not on a missing callback alone.

Deploy windows

The practical rule: if your normal deploy leaves the endpoint unreachable for longer than the window above, the provider's retry policy will not save you. With Stripe you have days of slack. With a Sume job webhook you have minutes, so schedule the sweeper to run right after every deploy.

Which window should drive your design

Design to the shortest window you depend on, not the longest one you read about. A team that integrates Stripe first tends to assume a webhook is effectively durable, then ports the same handler to a provider whose automatic retries end in minutes. The handler is fine; the missing piece is a recovery loop.

A reasonable pattern is two layers. The webhook is the fast path and finishes most jobs within seconds of completion. A poller over your own table of open jobs is the safety net, and it should ask Sume about each open job on a slow cadence with backoff. When the poller finds a terminal job you never heard about, it processes it through exactly the same code path as a webhook, keyed by job_id, so a late callback that finally arrives is a harmless duplicate.

Stripe also documents manual retries: Resend in the dashboard works for up to 15 days after the event and the CLI for up to 30 days. Sume's redeliver is per job and has no such cutoff stated in its docs, but do not rely on it as a plan. Keep the sweeper.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume