Sume run webhook: 10 s timeout, so queue before you call a slow model

Sume gives your webhook 10 seconds per attempt, up to 10 attempts. Verify the signature, store the event, return 2xx, then call any slow model. Python verifier.

4 min readSume
All posts

Short answer

A Sume run webhook has a 10 second budget per attempt, up to 10 attempts, and any 2xx counts as success. If your handler calls an LLM, such as Haiku 5.5, Astra or Mistral Large 4, before it answers, one slow completion can burn the whole attempt and Sume will retry. The docs say to record the event durably, return 2xx quickly, and process afterward.

This applies the same way whichever model you use, because the timeout is on Sume's delivery, not on the model.

Delivery rules

From the Run webhooks page, read 2026-10-08.

Delivery behavior (Sume docs)
PropertyValue
WhenOnce per run, on complete or fail
SuccessAny 2xx
RetriesMaximum 10 attempts, then status exhausted
Backoffmin(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h)
Timeout10 s per attempt
RedirectsNot followed; a 3xx is a failed attempt
Dedupe keyrequest_id, same on every retry
Body limit1,048,576 bytes; larger receipts are replaced by a result_url

Verifier

Sume signs <timestamp>.<raw_body> with HMAC-SHA256 and sends x-sume-webhook-signature: sume-v1=<hex> plus x-sume-webhook-timestamp. Reject timestamps outside your replay window; the docs suggest five minutes. This version refuses an empty secret.

import hashlib, hmac, time

def verify(raw_body: bytes, timestamp: str, signature: str, secret: str,
           tolerance: int = 300) -> bool:
    if not secret:
        raise ValueError("webhook secret is empty")
    try:
        ts = int(timestamp)
    except ValueError:
        return False
    if abs(int(time.time()) - ts) > tolerance:
        return False
    msg = str(ts).encode() + b"." + raw_body
    digest = hmac.new(secret.encode(), msg, hashlib.sha256).hexdigest()
    return hmac.compare_digest(f"sume-v1={digest}", signature)

Order of work

First verify, then write the event to a queue keyed by request_id, then return 200. A worker reads the queue and does the model call. The Format runs docs say the receipt is also available by polling, so a lost delivery never loses the result.

If the handler must call a model

When the post-processing really needs an LLM, for example summarizing the run's output.text into a message, do it after the response. A queue is the simple choice. Inside the worker, the model's own timeout is yours to set, since the vendor pages I read for Astra and Haiku 5.5 state none.

Test the retry path once. Make the handler return 500 and watch the delivery repeat with the same request_id; your queue key should absorb it. The backoff starts at 30 seconds and doubles, with jitter, up to an hour.

A short checklist for the handler, in order: read the raw body bytes before parsing, reject an empty secret at startup, verify the signature and the five-minute timestamp window, check the body against the one-megabyte limit, look up request_id and stop if it is already stored, write the event, and return a 2xx. Everything else, including any model call, runs after the response. Following that order means a slow model can never cause a retry, and a retry can never cause a duplicate effect.

Also log the delivery attempt number if the headers carry one, so a pattern of timeouts shows up in your own dashboards before Sume stops retrying.

If a delivery fails for good after ten attempts, fall back to polling the status URL, which stays available and holds the same final result as the event.

  • Dedupe on request_id.
  • A 3xx counts as failed, so point the URL at the final address.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume