Sume run webhook: 10 s timeout, so queue before you call a slow model
Sume gives your webhook 10 seconds per attempt, up to 10 attempts. Verify the signature, store the event, return 2xx, then call any slow model. Python verifier.

Short answer
A Sume run webhook has a 10 second budget per attempt, up to 10 attempts, and any 2xx counts as success. If your handler calls an LLM, such as Haiku 5.5, Astra or Mistral Large 4, before it answers, one slow completion can burn the whole attempt and Sume will retry. The docs say to record the event durably, return 2xx quickly, and process afterward.
This applies the same way whichever model you use, because the timeout is on Sume's delivery, not on the model.
Delivery rules
From the Run webhooks page, read 2026-10-08.
| Property | Value |
|---|---|
| When | Once per run, on complete or fail |
| Success | Any 2xx |
| Retries | Maximum 10 attempts, then status exhausted |
| Backoff | min(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h) |
| Timeout | 10 s per attempt |
| Redirects | Not followed; a 3xx is a failed attempt |
| Dedupe key | request_id, same on every retry |
| Body limit | 1,048,576 bytes; larger receipts are replaced by a result_url |
Verifier
Sume signs <timestamp>.<raw_body> with HMAC-SHA256 and sends x-sume-webhook-signature: sume-v1=<hex> plus x-sume-webhook-timestamp. Reject timestamps outside your replay window; the docs suggest five minutes. This version refuses an empty secret.
import hashlib, hmac, time
def verify(raw_body: bytes, timestamp: str, signature: str, secret: str,
tolerance: int = 300) -> bool:
if not secret:
raise ValueError("webhook secret is empty")
try:
ts = int(timestamp)
except ValueError:
return False
if abs(int(time.time()) - ts) > tolerance:
return False
msg = str(ts).encode() + b"." + raw_body
digest = hmac.new(secret.encode(), msg, hashlib.sha256).hexdigest()
return hmac.compare_digest(f"sume-v1={digest}", signature)Order of work
First verify, then write the event to a queue keyed by request_id, then return 200. A worker reads the queue and does the model call. The Format runs docs say the receipt is also available by polling, so a lost delivery never loses the result.
If the handler must call a model
When the post-processing really needs an LLM, for example summarizing the run's output.text into a message, do it after the response. A queue is the simple choice. Inside the worker, the model's own timeout is yours to set, since the vendor pages I read for Astra and Haiku 5.5 state none.
Test the retry path once. Make the handler return 500 and watch the delivery repeat with the same request_id; your queue key should absorb it. The backoff starts at 30 seconds and doubles, with jitter, up to an hour.
A short checklist for the handler, in order: read the raw body bytes before parsing, reject an empty secret at startup, verify the signature and the five-minute timestamp window, check the body against the one-megabyte limit, look up request_id and stop if it is already stored, write the event, and return a 2xx. Everything else, including any model call, runs after the response. Following that order means a slow model can never cause a retry, and a retry can never cause a duplicate effect.
Also log the delivery attempt number if the headers carry one, so a pattern of timeouts shows up in your own dashboards before Sume stops retrying.
If a delivery fails for good after ten attempts, fall back to polling the status URL, which stays available and holds the same final result as the event.
- Dedupe on
request_id. - A 3xx counts as failed, so point the URL at the final address.
Sources
Related posts
More in Developers
- Sume run webhook retries: ten attempts span about 3 hours
Ten delivery attempts with 30-second doubling backoff capped at one hour add up to 11,010 seconds before jitter. The timeline, and what to do after exhaustion.
- Sume run webhook says OK but output is null: handle degraded
A Sume run can complete, bill you, and still have output null. Branch on outcome, not status, and read artifacts and output_error before you retry.
- Sume STT: omit duration_seconds and you reserve one minute, not 570 s
The catalog says omitting duration_seconds reserves 1 minute; the maximum is 10 minutes ($0.10). A 570-second file lists at $0.095. Hold table, read 2026-10-08.
- Sume submit: request_id is the job id you poll (curl, jq)
After an async submit, data.request_id is the id for GET /v1/jobs/:id. error.request_id is for support tickets. A curl and jq check.
Written by Sume