waitForJob throws on a failed poll: resume by job id in TypeScript

Unlike waitForRun, waitForJob has no transient-failure budget: one failed status read after the client's retries throws. Wrap it and resume by job id.

5 min readSume
All posts

Does waitForJob survive a bad poll? Not past the client's own retries. In @sume-com/sdk@0.2.0, waitForJob throws a SumeJobRequestError the first time a status read comes back with an error, and waitForRun is the helper that absorbs up to six transient failures. The job keeps running and keeps billing when that happens, so the fix is a small wrapper that treats a failed read as a reason to look again, not a reason to give up.

The two layers matter. createSumeClient retries a GET on 408, 429, 5xx and transport errors up to maxRetries times (default 2) with backoff, honouring retry-after. Only when those retries are spent does the error reach waitForJob, which then throws. See Waiting for runs and jobs for the documented behavior of both helpers.

What is and is not covered

The table lays out where each failure lands. The job id is the handle: every path below leaves the job alive, so the same id can be passed to waitForJob again later.

Where a failed poll is handled in the Sume SDK (SDK source and docs, read 2026-10-04)
Failure on a status readClient retrywaitForJob result
429 or 5xxUp to 2 retries, backoff, retry-after honouredThrows SumeJobRequestError after retries
404 for the job idNone, not retryableThrows SumeJobRequestError at once
Job reaches failed or canceledNot an errorResolves with the job record
Deadline passes (20 min default)Not applicableThrows SumeJobTimeoutError

A resume wrapper

The wrapper below calls waitForJob in a loop under one overall deadline. A SumeJobRequestError with a 408, 429 or 5xx status, or a TypeError from fetch itself, pauses for a few seconds and calls it again with the same job id. Anything else, including a 404, is rethrown, because retrying a wrong id cannot help. A timeout is rethrown too, so you still get the lastStatus it carries.

import { createSumeClient, waitForJob, SumeJobRequestError } from "@sume-com/sdk";

const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });

function transient(status: number | undefined) {
  return status === undefined || status === 408 || status === 429 || status >= 500;
}

export async function waitForJobResumable(jobId: string, totalMs = 20 * 60_000) {
  const deadline = Date.now() + totalMs;
  for (let attempt = 0; ; attempt += 1) {
    try {
      return await waitForJob(jobId, {
        client,
        timeout: Math.max(5_000, deadline - Date.now()),
      });
    } catch (error) {
      const retry = error instanceof TypeError
        || (error instanceof SumeJobRequestError && transient(error.status));
      if (!retry || Date.now() >= deadline) throw error;
      const pause = Math.min(2_000 * 2 ** attempt, 30_000);
      console.warn("poll failed, resuming", { jobId, attempt });
      await new Promise((resolve) => setTimeout(resolve, pause));
    }
  }
}

Rules to keep with it

  • Persist the job id the moment the submit returns it, before any waiting. A process restart then costs a re-read, not a lost job.
  • Submit with an Idempotency-Key so a retried submit cannot create a second job. The wrapper retries reads only, never the submit.
  • A resolved job is not always a successful one. Check status for completed and read error otherwise.
  • For long video work, prefer a webhook and keep polling as the fallback. The webhooks guide covers delivery.

Related posts

More in Developers

All Developers posts

Written by Sume