Graceful shutdown for a Node webhook receiver: 503 while draining

On SIGTERM, refuse new Sume webhooks with 503, finish the in-flight write, then exit. Sume retries non-2xx responses 10 times, so deploys lose nothing.

5 min readSume
All posts

On SIGTERM, flip a draining flag, answer every new webhook with 503 and a retry-after header, wait for the requests that are already in flight to finish their durable write, and then exit. Sume treats a non-2xx response as a failed attempt and retries, so a 503 during a deploy costs one attempt of ten and loses nothing. A 200 that you never stored is the one answer that loses an event.

The delivery rules in Sume's webhook docs make this safe: up to 10 attempts in total, a fixed delay between them (30 seconds by default), a 10 second timeout on each, and the advice to return a 2xx only after you have stored the event durably.

A receiver that drains

The server below stops taking work first, finishes what it has, and then exits. closeIdleConnections() matters because keep-alive connections would otherwise hold server.close() open. Replace store with your verify-and-write code.

import http from "node:http";
let draining = false, inflight = 0;
const store = async (raw) => { /* verify the signature, then write job_id durably */ };

const server = http.createServer(async (req, res) => {
  if (draining) {
    res.writeHead(503, { "retry-after": "30", connection: "close" }).end();
    return;
  }
  inflight++;
  try {
    const chunks = [];
    for await (const c of req) chunks.push(c);
    await store(Buffer.concat(chunks).toString("utf8"));
    res.writeHead(204).end();
  } catch { res.writeHead(500).end(); } finally {
    inflight--;
  }
});

process.on("SIGTERM", () => {
  draining = true;
  server.close(() => process.exit(0));
  server.closeIdleConnections();
  setTimeout(() => process.exit(inflight ? 1 : 0), 8000).unref();
});

server.listen(8080);

What each answer does during a deploy

The exit timer is 8 seconds on purpose. Sume gives each attempt 10 seconds, so a request that is still running after that is already lost on Sume's side.

Receiver answers while draining and what Sume does (Sume docs, read 2026-10-04)
MomentYour answerSume behavior
Before SIGTERM204 after the durable writeDelivery done
SIGTERM, new request arrives503 with retry-afterCounts as a failed attempt, retries later
SIGTERM, request already writing204 when the write finishesDelivery done
Process killed mid-writeConnection resetRetries; keep job_id as the dedupe key
All 10 attempts failNothingJob is still terminal; Redeliver or poll recovers it

When the outage is longer than a deploy

Two safety nets cover a longer outage. First, POST /v1/jobs/{job_id}/webhook/redeliver (with jobs:write) re-sends the real terminal event with a fresh timestamp and signature, even after the automatic attempts are used up. Second, status_url polling shows the true state of any job you expected to hear about. Because both can repeat an event you already handled, key your write on job_id.

Caveats

  • A 503 is a signal to retry. A 4xx can also be retried by Sume, so use 5xx for "try again later" and 401 only for a bad signature.
  • In Kubernetes, give the pod a termination grace period longer than the 8 second timer, and stop routing traffic to it before SIGTERM if your platform lets you.
  • Do not use this pattern to hide a slow handler. Sume's 10 second budget applies to every attempt, including the ones that succeed.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume