OpenAI retries webhooks 72 hours; Sume stops after ten attempts

OpenAI retries a failing webhook for up to 72 hours. Sume stops after ten attempts, about 5 minutes for jobs and 3 hours for runs, so keep a status poll.

5 min readSume
All posts

OpenAI says it retries webhook deliveries for up to 72 hours with exponential backoff. Sume stops after ten attempts. At the default 30-second spacing for job webhooks that is about four and a half minutes of gaps, and for Format run webhooks it is roughly three hours. If your receiver can be down longer than that, you need a poll or Redeliver as the backstop.

The numbers

OpenAI's webhooks guide also says to respond within a few seconds with a 2xx, to treat duplicates as possible, and to use the webhook-id header as the idempotency key. Sume's equivalent is job_id or request_id. The retry windows are what differ.

The Sume totals below are my arithmetic from the documented schedule, ignoring jitter and the 10-second timeout of each attempt, and reading the run formula as the delay after each failed attempt.

Retry windows (OpenAI read 2026-10-02, Sume docs)
WebhookScheduleTotal window
OpenAIExponential backoffUp to 72 hours
Sume job webhooksFixed 30s by default, 10 attemptsAbout 4.5 minutes of spacing
Sume Format run webhooks30s x 2^(attempt-1), capped at 1h, 10 attemptsRoughly 3 hours of spacing

What happens when the window closes

Sume's docs say that a delivery which runs out of attempts leaves a failed delivery and a job or run that still reached its real terminal state. A run's webhook_delivery.status becomes exhausted. Nothing is lost: the result is still on result_url.

  • Poll status_url on a slow cadence for every call that has a webhook, and stop on the terminal boolean.
  • Check webhook_delivery.status on Format runs for exhausted.
  • Use Redeliver after a long outage. It works after the ten attempts and does not use one of them.
  • Respect next_poll_after_seconds when polling, and the read budget for your plan.

Which one to build around

If your receiver has a good uptime record, the webhook alone is fine, with a daily sweep for anything stuck. If it runs on a host that is often rebooted or rate-limited, treat the webhook as a hint and make polling the source of truth. Sume's docs call delivery an optimization, never the only recovery path, which is the right way to design for the shorter window.

Alerting on the gap

Add two alerts. The first fires when a call with a webhook has been non-terminal for longer than you expect, which catches both slow jobs and missing deliveries. The second fires when a run shows webhook_delivery.status of exhausted, which tells you your receiver was down for the whole window.

When either fires, read the job or run by id, process the result with the same code path your webhook handler uses, and mark it done in your dedupe table. That way a recovered item and a late delivery cannot both be processed. Redeliver is available once the receiver is healthy again, but it is optional when you have already read the result.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume