Return 503 with Retry-After in a deploy: Sume run webhooks honour it

Sume run webhooks honour Retry-After on 429 and 503 up to 1 hour. Job webhooks use a fixed 30s spacing, so keep the drain short and the fallback poll on.

5 min readSume
All posts

If your receiver is going down for a deploy, answer 503 with a Retry-After header. Sume's run webhooks page says it honours Retry-After on 429 and 503, and the delay is capped at 1 hour. Job webhooks work differently: the docs describe a fixed delay of 30 seconds between attempts and do not mention Retry-After, so keep the outage short for those.

The header, from MDN

MDN describes Retry-After as having two forms: an HTTP date, or a non-negative number of seconds to wait after the response is received. It applies to 503, where it says how long the service is expected to be unavailable, and to 429. MDN also notes that support on clients and servers is inconsistent, so do not assume every sender honours it.

Retry behaviour by webhook family (Sume docs, MDN read 2026-10-02)
FamilySpacingRetry-After
Format run webhooksmin(max(30s x 2^(attempt-1) with jitter, Retry-After), 1h)Honoured on 429 and 503
Job webhooksFixed delay, 30s by defaultNot mentioned in the docs

A drain handler

Flip a flag at the start of the deploy, return 503 with a delay a bit longer than the restart, and flip it back. This Node 18+ sketch shows only the drain branch. In a real receiver the signature check and the durable write come first, as in the other posts in this series.

import http from "node:http";

let draining = false;
process.on("SIGTERM", () => { draining = true; });

const server = http.createServer((req, res) => {
  if (draining) {
    res.writeHead(503, { "Retry-After": "120" });
    res.end();
    return;
  }
  // verify the signature, store the event, then answer
  res.writeHead(200);
  res.end();
});

server.listen(3000);

How long is too long

There are ten attempts in total on both families, and each attempt gets 10 seconds. With job webhooks at the default 30-second spacing, the spacing alone adds up to about four and a half minutes across the nine gaps, so a longer outage can exhaust the budget. The job still finishes. Recover with GET /v1/jobs/:id on status_url, or use Redeliver on the job once you are back, which does not use one of the automatic ten.

For a rolling deploy with at least one healthy instance, you do not need any of this. Return 503 only when no instance can take the request.

Testing the drain

Run the sketch locally, send it SIGTERM, then POST to it and confirm you see a 503 with Retry-After: 120. A second POST should keep seeing the same answer until the process exits. Then restart it and confirm the next POST goes through the normal path.

In production, pair the drain with a readiness check so your load balancer stops sending new traffic before the process stops. A Retry-After shorter than your restart time just wastes an attempt, and one that is much longer delays a delivery that could have landed. Pick a value slightly longer than your slowest restart, and never above the 1 hour cap Sume applies on run webhooks.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume