Agent Completion webhook retries: 10 attempts over about 3 hours
Sume tries an agent.run.terminal webhook up to 10 times. By the documented formula the nine waits add up to about 3 h 3 min, before jitter and Retry-After.

If your endpoint keeps failing, Sume stops after 10 attempts, and by the documented backoff formula the nine waits between them add up to roughly 3 hours 3 minutes. That is before jitter and before any Retry-After header your endpoint sends, so treat it as an estimate. After the tenth attempt the run's webhook_delivery.status becomes exhausted, but the run itself stays completed or failed, and you can still fetch it from result_url.
The page's formula is min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h). Each attempt has a 10-second timeout, a 2xx is success, and a 3xx counts as a failed attempt because Sume does not follow redirects.
The schedule, computed
I read the formula as the wait after attempt k, with the cap at 3,600 seconds. Jitter and Retry-After are left out.
| After attempt | 30 x 2^(k-1) s | Capped wait (s) |
|---|---|---|
| 1 | 30 | 30 |
| 2 | 60 | 60 |
| 3 | 120 | 120 |
| 4 | 240 | 240 |
| 5 | 480 | 480 |
| 6 | 960 | 960 |
| 7 | 1,920 | 1,920 |
| 8 | 3,840 | 3,600 |
| 9 | 7,680 | 3,600 |
def waits(attempts=10, base=30, cap=3600):
return [min(base * 2 ** (k - 1), cap) for k in range(1, attempts)]
total = sum(waits())
print(waits())
print(total, divmod(total, 3600), total / 3600)
# 11010 seconds = 3 h 3 min 30 sWhat to build around it
Sum of the nine waits: 30 + 60 + 120 + 240 + 480 + 960 + 1,920 + 3,600 + 3,600 = 11,010 seconds, or 3 hours 3 minutes 30 seconds.
- Record the event durably, return
2xxquickly, and process afterward. A slow handler burns the 10-second budget and triggers another attempt. - Dedupe on the envelope's
request_id, which equals the run id and is stable across retries. - A
canceledrun and askippedrun send no webhook at all, so do not wait for one. - If deliveries exhaust, use Redeliver on the delivery row, or fetch the receipt from
result_url. Redeliver is documented for Format runs; for Agent Completions pollstatus_url.
Signature check, briefly
Sume signs <timestamp>.<raw_body> with HMAC-SHA256 and sends x-sume-webhook-signature: sume-v1=<hex>. Verify against the raw body before parsing, reject timestamps outside a window (five minutes is the suggested default), and refuse to run at all if your secret is empty, since an empty secret makes every signature forgeable.
Planning around the window
About three hours is the span in which your endpoint can recover and still get an automatic delivery. If your deploys or outages run longer, do not rely on the webhook alone: keep status_url polling as a backup, which the docs explicitly allow. A reconciliation job that lists recent runs and fetches any that never reported is cheap insurance, and it also covers the canceled and skipped cases that never send a webhook.
Note that a failed delivery never changes the run. The generation has already been billed against the cap, so a missed webhook costs you latency, not money.
Sources
Related posts
More in Developers
- AI video API billing units: per clip, per second, credits or tokens
Luma bills per generation, LTX per second, Runway and Vidu in credits, Google in tokens, MiniMax per second plus inputs, Sume in USD. Worked examples.
- Video fields Sume rejects: size, seed, provider.options, audio off
Sume's /v1/videos refuses size, seed and non-empty provider.options on every model, and Omni refuses generate_audio false. What to send in their place.
- AI voiceover too loud: three Sume gain knobs and what each one costs
TTS generation_config.volume (0.5-2), Timeline audio.gain_db (-60 to 12) and soundtrack.duck_db (0-20). Which to change, and which means paying for new audio.
- Alibaba Wan 3.0 Model Studio request to a Sume /v1/videos body
Map a Model Studio wan3.0-video call (input.media, parameters, X-DashScope-Async) to POST /v1/videos with model wan-3.0. Field by field.
Written by Sume