Test a Sume poll loop without waiting: inject sleep, assert delays

Unit test a job poll loop in milliseconds by injecting the fetch and the sleep. Assert that next_poll_after_seconds is obeyed and the 20-minute deadline holds.

5 min readSume
All posts

To unit test a Sume poll loop without waiting for real time, pass the loop two functions it calls instead of importing them: a getStatus that returns scripted envelopes and a sleep that records the number of seconds it was asked for. The test then asserts the list of delays, which is the actual contract: obey next_poll_after_seconds when it is present, back off when it is missing, and stop at a client-side deadline. A 20-minute video wait finishes in under a millisecond, with no timers, no mocks of the clock and no flaky retries in CI. This follows the loop in Sume jobs and results.

What does the loop need to prove?

Sume says to poll status_url until terminal is true, to use next_poll_after_seconds when present and an exponential backoff when not, and to keep the overall deadline on the client. It also says a client timeout does not cancel the job: the job keeps running and billing. So the loop has three behaviours worth a test, and a time-based test cannot reach the deadline one in a reasonable runtime.

Fake timers are the usual answer, but they couple the test to a framework's clock. Passing the sleep in is plainer: the production call is sleep = s => new Promise(r => setTimeout(r, s * 1000)), and the test passes a function that pushes to an array and returns at once.

What does the test look like?

const assert = require("node:assert/strict");
async function waitForJob(getStatus, sleep, { deadlineSeconds = 1200 } = {}) {
  let waited = 0, backoff = 2;
  for (;;) {
    const s = await getStatus();
    if (s.terminal) return s;
    const delay = s.next_poll_after_seconds ?? backoff;
    backoff = Math.min(backoff * 2, 30);
    if (waited + delay > deadlineSeconds) throw new Error("deadline: " + s.request_id);
    waited += delay;
    await sleep(delay);
  }
}
const script = (list) => async () => list.shift();
(async () => {
  const delays = [];
  const sleep = async (s) => delays.push(s);
  const id = { request_id: "job_1" };
  const wait = { ...id, terminal: false };
  await waitForJob(script([
    { ...wait, next_poll_after_seconds: 5 }, wait, wait, { ...id, terminal: true },
  ]), sleep);
  assert.deepEqual(delays, [5, 4, 8]);
  const never = async () => ({ ...wait, next_poll_after_seconds: 600 });
  await assert.rejects(waitForJob(never, sleep, { deadlineSeconds: 1200 }), /deadline/);
  console.log("ok", delays);
})();

What do the delays prove?

Read the assertion [5, 4, 8] closely. The first envelope carries a hint, so the loop sleeps 5. The next two have no hint, so the backoff takes over. It doubles on every poll, including the hinted one, so the first unhinted sleep is 4 and the next is 8; the hint did not reset it, which is a choice you may want to flip. The test pins whatever you decide, and that is its value: changing the policy later fails loudly instead of quietly polling every second.

The deadline case throws with the job id in the message on purpose. Since a client-side timeout leaves the job running, the error is the only place you still hold the id, and the right follow-up is to store it, then read the status later or cancel it explicitly. The same fact explains why a deadline is not the same thing as a cancel. A cancel after generation has started comes back 409 job_generation_already_started, so a deadline handler should be ready to see that and leave the job to finish instead of treating the cancel failure as a crash. In both cases the record you wrote at submit time is what lets you find the job again.

Cases to cover in the poll test, from Sume jobs and results (read 2026-10-06)
CaseScripted inputAssert
Hint presentnext_poll_after_seconds: 5sleep(5)
Hint missingno hint fieldBackoff grows, capped
Deadlinehint of 600, limit 1200Throws with request_id
Terminal first pollterminal: trueNo sleep at all

What else can the seam test?

Two refinements are worth the extra lines. Add jitter to the backoff in production and make the jitter injectable too, as a random argument, so the test stays exact. And make the deadline a clock check rather than a running total when you poll a very long job, because the time spent inside getStatus counts against the 20-minute budget as well, and a slow network can make the total drift past what you meant.

The same seam gives you a place to put progress reporting. Pass an onStatus callback, call it once per poll, and the test can assert that a UI would have seen queued, then processing, then completed, without any rendering code in the way.

Where does the real fetch get tested?

Keep the network out of this test entirely. A separate, thin test can check that the real getStatus sends one credential header and reads the envelope; the loop test should need neither a key nor a server.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume