How long to wait for a Sume job webhook before you start polling

Job webhooks retry 10 times, 30 seconds apart, with a 10-second timeout each. Start your poll fallback at about 6 minutes, and here is the arithmetic.

4 min readSume
All posts

Wait about six minutes after a job's terminal time before you treat a missing job webhook as lost, then read status_url. Sume sends up to 10 attempts, spaced by a fixed delay that is 30 seconds by default, with a 10-second timeout on each attempt. Nine gaps of 30 seconds is 270 seconds; add up to 10 seconds of timeout per attempt and the longest automatic window is under about 6 minutes and 10 seconds. After that, no further automatic attempt will come.

The arithmetic

These numbers are the documented defaults. Your deployment's delay is whatever the docs' 30-second default gives you unless you were told otherwise, so treat the table as the default case.

Job webhook delivery window at the documented defaults (read 2026-10-05)
QuantityValueSource of the number
Attempts10 totalWebhooks page, delivery behavior
Gap between attempts30 s fixed, not exponentialWebhooks page
Gaps in a full run910 attempts have 9 gaps
Time spent in gaps270 s (4.5 minutes)9 x 30 s
Worst-case time in attempts100 s10 x 10 s timeout
Longest automatic window370 s (about 6 min 10 s)270 s + 100 s

What to do at each point

Before the window closes, a missing webhook can still arrive, so a poll in that period is only a cross-check. Polling does not interfere with delivery. After the window closes, a job that is terminal on status_url but silent on your endpoint means the delivery failed ten times: your endpoint was down, slow past 10 seconds, or answering non-2xx.

At that point, fix the endpoint and use the documented redeliver action: POST /v1/jobs/{job_id}/webhook/redeliver with jobs:write, which re-POSTs the real terminal event with a fresh timestamp and signature. It still works after the automatic attempts are used up and does not consume one of the ten.

A simple sweep

Keep a table of open chain steps with the time each job was submitted. On a timer, select steps still open that were submitted longer ago than your longest expected job plus the 370-second window, and read each status_url. A step that is terminal gets processed through the same handler the webhook uses, deduped on job_id, so a late webhook and the sweep cannot both run the next step.

For jobs that are still queued or processing, a poll sweep is not a failure signal at all. A 30-second video can wait in the queue when the workspace is at its concurrency limit, and queued is a normal accepted state.

Runs are different

Do not apply these numbers to Action, Format or Agent Completion runs. Their *.run.terminal webhooks use the same 10-attempt cap on a different schedule, and the run webhooks page is the authority for that schedule. A canceled run sends no webhook at all, so for runs you poll status_url after cancel.

A hybrid that stays cheap

The practical answer is to use both. Register the webhook as the main signal and run one slow poll as a safety net. The webhook handles the usual case with no polling traffic at all. The poll catches the rare case in which a delivery never arrives, such as a receiver outage that outlasts the 10 attempts.

Use the delivery numbers from the docs to set the net: 10 attempts with a 30 second gap is about 4.5 minutes of automatic retrying after the first try. A single status read after that window, and then at a long interval, is enough. If you find that a delivery was missed, POST /v1/jobs/{job_id}/webhook/redeliver re-sends the terminal event with a fresh signature.

Reads are cheap: the read bucket is 40 times the write bucket on the shipped default. A slow safety poll costs almost nothing against it.

  • Webhook first, one slow poll second.
  • Start the net after the retry window.
  • Redeliver by hand when a delivery was lost.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume