Move from polling to Sume webhooks in three steps, keeping the poll

Add webhook_url to submits, verify sume-v1 signatures, dedupe on job_id, and keep a slow poll for canceled and skipped runs. Roll it out one route at a time.

6 min readSume
All posts

Do it in three steps and never remove the poll. First, stand up a receiver that verifies the sume-v1 signature and answers 2xx quickly. Second, pass webhook_url on new submits, which alone selects webhook mode. Third, slow the poll down to a safety net instead of deleting it. The docs describe a webhook as the alternative to a poll loop, and they say you still get status_url and result_url and can still poll.

The reason to move is load and latency. A poll loop spends read budget on every check, even though most checks find nothing new, and a webhook sends one request when something actually happened. Reads are generous, with their own bucket per key, but a fleet of watchers still adds up, and each watcher also adds code that you have to keep alive.

Step 1: a receiver that is safe to deploy

Build the receiver first, and prove it with the dummy event from POST /v1/webhooks/test-deliveries. It must read the raw body, check the timestamp inside the 300 second window, and compare the HMAC over <timestamp>.<raw_body> against the sume-v1=<hex> header. Refuse to start when the secret is empty. Answer 204 within the 10 second delivery timeout, and do the real work after the response.

Treat the receiver as a normal public endpoint. Put it behind HTTPS, keep it small, and let it write to a queue or a table and return. Reading the signing secret needs account:read, and you can also copy it from the Webhooks tab in the dashboard. Keep it in the environment variable SUME_COM_WEBHOOK_SIGNING_SECRET, the same name the docs and the SDK use.

Step 2: turn it on for one route

Register a public HTTPS URL, at most 2048 characters, on each submit. Sume rejects http, localhost and private addresses with a 400 at submit time, and checks them again at delivery time. It does not follow redirects, so register the final URL. Start with one low risk route, such as thumbnails, and move the others after a week of clean deliveries.

Mixed fleets are normal during a rollout. Jobs submitted before the change have no webhook and still need their poll. Jobs submitted after it have both. Do not switch the poll off by submit date alone. Switch it off per route, and only after you have seen the delivery status of real jobs reach delivered.

What a webhook covers and what it does not

Compare what the two paths give you before you trust a webhook alone.

Webhook coverage compared with polling (read 2026-10-05)
CaseWebhookPoll
Job completed or failedjob.completed, job.failedStatus reaches a terminal value
Job canceledjob.canceledStatus is canceled
Run completed or failedformat.run.terminal and its siblingsRun status is terminal
Run canceled or skippedNo webhook is sentOnly a poll shows it
Receipt over 1 MiBpayload is null, with a result_urlFull receipt from the poll route
Endpoint was downTen attempts, then exhaustedAlways available

Step 3: demote the poll to a safety net

Keep the poll, but make it slow. A job or run with no webhook after a generous wait is the case to catch. Read its status once, and if it is terminal, handle it with the same function that your webhook route calls. If the webhook arrived but your database write failed, POST /v1/jobs/{id}/webhook/redeliver with jobs:write sends it again, and the Format run route needs formats:write. A redelivery uses a fresh timestamp and signature and does not use one of the 10 attempts.

The same rule helps when you replay history. If you ever load old events from a log into the receiver, the dedupe key stops them from doing damage, and the no-backwards rule stops an old event from overwriting a newer state.

One idempotent write path

Both paths must write through one function, and that function must be idempotent. Dedupe on job_id or request_id, and never move a row backwards, because a late webhook can arrive after the poll already stored the final state. The sample below is that function in plain Python with SQLite, and it runs as is.

Also decide who gets paged. A job whose delivery is exhausted still finished fine on the Sume side, so the right response is a redelivery or a poll, not a resubmit of the paid request.

import sqlite3

db = sqlite3.connect(":memory:")
db.execute("create table jobs (job_id text primary key, status text)")

def apply_terminal(job_id: str, status: str) -> bool:
    cur = db.execute(
        "insert into jobs values (?, ?) on conflict(job_id) do nothing",
        (job_id, status),
    )
    db.commit()
    return cur.rowcount == 1  # True only the first time

print(apply_terminal("job_demo", "completed"))  # webhook
print(apply_terminal("job_demo", "completed"))  # poll safety net

Watch the delivery state

Watch the delivery state while you roll out. The delivery statuses are pending, delivering, delivered, retrying, failed and exhausted, and the job record shows them with the attempt count. A rise in retrying means your receiver is slow or erroring. Fix that before you widen the rollout, and keep the 10 attempt budget in mind: job webhooks use a fixed 30 second spacing, so the window for a down receiver is only a few minutes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume