Cancel queued Sume jobs on SIGTERM during a deploy

On shutdown, cancel jobs you no longer need before they start and leave started ones alone. A 21-line Node handler using POST /v1/jobs/{id}/cancel.

4 min readSume
All posts

When a worker that submitted Sume jobs gets SIGTERM during a deploy, call POST /v1/jobs/{id}/cancel for the jobs it still tracks. Cancelling succeeds only before generation starts, so queued jobs are cancelled and their reservations released, while a job that has started answers 409 job_generation_already_started and runs to the end. Treat that 409 as a normal outcome and persist those ids so the next process can keep watching.

This matters most for a worker with a big fan-out on a small plan. Free allows 6 accepted jobs and Pro 24, per the generation admission page, and an abandoned queue of jobs nobody will read still occupies that capacity until it runs, and bills when it completes.

What cancel does and does not do

Cancel semantics, read 2026-10-02 from docs.sume.com
Job stateCancel resultBilling
queuedCanceled; the job becomes terminalReservation released or refunded where applicable
processing, generation started409 job_generation_already_started, details.cancelable falseCompletes or fails normally
already canceledIdempotent; returns the same canceled jobNone
completedNot cancelableAlready captured

The handler

Keep a Set of ids this process submitted and has not seen finish; delete from it when a job reaches a terminal state. Promise.allSettled makes one failing cancel not hide the others, and the function returns a pair per id for logging.

const BASE = process.env.SUME_BASE_URL ?? "https://api.sume.com";
const headers = { Authorization: `Bearer ${process.env.SUME_API_KEY}` };
export const inFlight = new Set(); // job ids this process submitted and has not finished

export async function cancelPending(ids = [...inFlight]) {
  const results = await Promise.allSettled(ids.map(async (id) => {
    const res = await fetch(`${BASE}/v1/jobs/${id}/cancel`, { method: "POST", headers });
    if (res.ok) return "canceled";
    const body = await res.json().catch(() => ({}));
    if (body.error?.code === "job_generation_already_started") return "running"; // will finish and bill
    throw new Error(`${id}: ${res.status} ${body.error?.code ?? ""}`);
  }));
  return ids.map((id, i) => [id, results[i].status === "fulfilled" ? results[i].value : results[i].reason.message]);
}

export function cancelOnShutdown(exit = process.exit) {
  process.once("SIGTERM", async () => {
    for (const [id, outcome] of await cancelPending()) console.log(id, outcome);
    exit(0); // running jobs keep going: store their ids so a restart can resume watching
  });
}

Wiring it in

The decision to cancel is a product decision. A rolling deploy that cancels every in-flight job turns each deploy into a customer-visible failure.

  • Add the job id to inFlight straight after a successful submit, and persist it to your database too.
  • Remove it when status reads completed, failed or canceled.
  • Give the process a termination grace period longer than a few cancel calls. Container platforms usually kill after a fixed timeout.
  • Do not cancel jobs a user is still waiting for just because your process is restarting; hand them over by id instead.

Limits

I tested it on Node 22 against a mock returning a cancelled job, the documented 409, and a 500: the output was canceled, running and an error line for the three ids. SIGKILL and crashes skip handlers, so this reduces waste but cannot replace recovery by id. Cancelling is a write, so it spends the write budget, 120 per minute on Free up to 1200 on Scale; with at most 120 accepted jobs on any self-serve plan, one sweep fits inside it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume