Abort waitForJob on SIGTERM, then resume from the stored job id
On a deploy, abort waitForJob with an AbortController on SIGTERM, keep the Sume job id, and resume the wait after restart. Aborting stops the wait, not the job.

To stop a Sume job wait cleanly on deploy, create an AbortController, pass its signal to waitForJob, and call abort() from your SIGTERM handler. The SDK documents signal as aborting the wait and the in-flight request. It does not cancel the job: the job keeps running and keeps spending on Sume's side, so persist the job id before you wait and resume with the same id after restart.
Node's process docs list SIGTERM as a signal a process can listen for with process.on (read 2026-10-04). Container platforms typically send it before a hard kill, so you have a short window to record where you were.
What does the pattern look like?
Store the id first, then wait. On start, look for an unfinished id and wait on that instead of submitting again.
import { createSumeClient, waitForJob } from "@sume-com/sdk";
import { readFileSync, writeFileSync, existsSync } from "node:fs";
const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });
const STATE = "job-id.txt"; // use your database in production
const ac = new AbortController();
process.on("SIGTERM", () => ac.abort(new Error("shutting down")));
async function main() {
if (!existsSync(STATE)) {
throw new Error("no stored job id; submit first and write it to " + STATE);
}
const id = readFileSync(STATE, "utf8").trim();
try {
const job = await waitForJob(id, { client, signal: ac.signal });
console.log(id, job.status);
} catch (err) {
if (ac.signal.aborted) {
console.log("stopped waiting; job", id, "is still running");
return;
}
throw err;
}
}
main().catch((e) => {
console.error(e);
process.exitCode = 1;
});What does an abort change on Sume's side?
| Action | Wait loop | Job on Sume's side |
|---|---|---|
| signal.abort() | Rejects with the signal's reason, in-flight read aborted | Keeps running and spending |
| Client timeout elapses | Throws SumeJobTimeoutError | Keeps running |
| Process killed | Gone | Keeps running; the id is your only handle |
| Cancel request | Next status read shows canceled | Moves toward canceled; confirm by polling |
How do you resume without double spend?
Resume by waiting on the stored id, not by submitting again. If you must resubmit, send the same Idempotency-Key as the first call: Sume replays the original job rather than starting a second one, as Jobs and results describes. The key has to be derived from your own record, not from Date.now().
What else belongs in the shutdown path?
- Write the job id to durable storage before the first status read, not after.
- Treat the abort rejection as a normal exit, not an error to alert on.
- If the job should stop with the process, send an explicit cancel; aborting the wait will not do it.
- For minutes-long video, a webhook removes the need for a process to be alive at all.
Sources
Related posts
More in Developers
- waitForJob throws on a failed poll: resume by job id in TypeScript
Unlike waitForRun, waitForJob has no transient-failure budget: one failed status read after the client's retries throws. Wrap it and resume by job id.
- waitForRun maxTransientFailures: how many bad polls it absorbs
waitForRun absorbs 6 consecutive 429, 5xx or network read failures before it throws. How the streak resets, backoff works, and onTransientError fits in.
- Wan 3.0 at 30 fps and 2 to 30 s: the Sume model id
Model Studio lists Wan 3.0 at 480P to 1080P, 2 to 30 s, 30 fps. On Sume the id is wan-3.0, 2 to 30 s; first and last frames go in frame_images.
- Wan 3.0 has no fine-tuning or batch mode; what Sume requests allow
Alibaba's Wan 3.0 docs list batch inference, fine-tuning and function calling as unsupported. Sume adds its own limit: non-empty provider.options is a 400.
Written by Sume