Ack a Sume webhook in 10 s, then queue it in SQS

Sume gives each webhook attempt 10 seconds, and Stripe advises a quick 2xx before heavy work. Verify, enqueue to SQS, return 204, size the visibility timeout.

6 min readSume
All posts

Short answer

Verify the signature, put the raw event on a queue, and return a 2xx within the 10 seconds Sume allows for each attempt. Stripe gives the same advice: return a 2xx quickly before complex logic and process asynchronously through a queue. With Amazon SQS, set the visibility timeout above your worker's longest run, because the default is 30 seconds and delivery is at least once.

The three clocks

A webhook path has three time limits and mixing them up is a common source of duplicate work. Sume's delivery worker waits 10 seconds per attempt and makes up to 10 attempts. SQS hides a received message for its visibility timeout. Your worker needs enough time to finish inside that window.

Redirects count as failures for Sume deliveries, so the endpoint should answer directly with a 2xx rather than bounce through a gateway redirect.

Limits that shape a webhook queue (read 2026-10-03)
ItemValueSource
Sume attempts per deliveryUp to 10Sume run webhooks docs
Sume timeout per attempt10 secondsSume run webhooks docs
Stripe adviceReturn 2xx quickly; process via a queueStripe webhooks docs
SQS default visibility timeout30 secondsAWS SQS docs
SQS maximum visibility timeout12 hours from first receiveAWS SQS docs
SQS deliveryAt least onceAWS SQS docs

The handler

Read the raw body first and verify the sume-v1 signature before anything else. If verification fails, return 401 and do not enqueue. If it passes, send the body and the delivery headers to the queue as one message, then return 204. Do not parse, call your database or call another API before you answer.

Keep the message small. Sume's run webhooks send a receipt over 1 MiB with a null payload, a payload_too_large error and a result_url, so your handler should be ready to fetch the result later by URL instead of expecting the full body inline.

The worker and the visibility timeout

SQS can deliver a message more than once, and a worker that outlives the visibility timeout will see the message handed to a second consumer. Set the timeout above the longest realistic processing time. AWS notes that the maximum is 12 hours from first receipt and that extending visibility does not reset that clock.

Make the worker idempotent regardless. For run webhooks, use the envelope's request_id (it equals the run id and is stable across retries) as the dedupe key and order deliveries by created_at; for job webhooks, use job_id as the idempotency key. A duplicate then becomes a no-op. For persistent failures, configure a dead-letter queue so a poison message stops cycling.

Why not process inline

Inline work makes every slow dependency your webhook's problem. If your database stalls for 12 seconds, Sume sees a timeout, marks the attempt failed and retries, and your system ends up processing the same event twice while still behind. Acknowledging first moves that risk into a queue that you can size, monitor and replay.

One caveat: a successful ack means you accepted the event, not that you handled it. Track queue age and dead-letter depth, and keep a slow poll of job or run status as a backup for anything the queue loses. Job webhooks retry on a fixed 30 second spacing and run webhooks back off exponentially; see the two retry schedules.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume