waitForRun maxTransientFailures: how many bad polls it absorbs

waitForRun absorbs 6 consecutive 429, 5xx or network read failures before it throws. How the streak resets, backoff works, and onTransientError fits in.

5 min readSume
All posts

waitForRun from @sume-com/sdk@0.2.0 tolerates 6 consecutive transient read failures by default (maxTransientFailures: 6). A transient failure is a status read that fails with a retryable error: a 429, a 5xx, or a transport error. The seventh in a row throws a SumeRunRequestError. Any clean read resets the count to zero. A non-retryable error, such as a 404 for the wrong run id, throws immediately.

The reason is in the SDK's own comments: a failed read is not a failed run. The run is still executing and still spending, and throwing from the poll loop would hand your code an exception while the meter keeps running and you no longer hold a handle to the result. See the Sume SDK run helpers for the full option table.

What counts, and what resets the streak

The SDK builds a typed error from each failed read and checks its retryable flag. That flag comes from the error envelope when the server sent one, and from the HTTP status otherwise (408, 429 and 5xx are retryable). The counter only moves on retryable failures, and only a successful status read clears it, so alternating failure and success cannot spin forever on a budget that never refills.

The same budget applies once the run is terminal. The final receipt read is retried with the same limit, so a 429 on that last call does not throw away a finished result you already paid for.

How waitForRun treats a failed status read (Sume SDK source, read 2026-10-04)
FailureRetryableBehavior
429 with retry-afterYesWaits the server's window (capped at 60 s), jittered
5xx or network errorYesWaits poll interval x 2 per streak step, capped at 30 s, jittered
404 or 403NoThrows SumeRunRequestError immediately
Seventh retryable failure in a rowYesThrows SumeRunRequestError
Backoff would pass the timeoutYesThrows SumeRunTimeoutError

Watching absorbed failures with onTransientError

onTransientError(error, attempt) fires each time a failure is absorbed instead of thrown. The run is still executing when it fires, so use it for a log line or a metric, not for cancel logic. With timeline: true on a Format run, a failed timeline read is reported through the same callback and the previous timeline is kept.

import { createSumeClient, waitForRun } from "@sume-com/sdk";

const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });

const run = await waitForRun(process.argv[2]!, {
  client,
  family: "format",
  maxTransientFailures: 6,
  onTransientError: (error, attempt) => {
    console.warn(JSON.stringify({
      event: "sume_poll_transient",
      attempt,
      status: error.status,
      code: error.code,
      request_id: error.requestId,
    }));
  },
  onStatus: (status) => console.log("status", status),
});

console.log(run.status);

When to change the number

Set maxTransientFailures: 0 if you want the first hiccup to surface, for example in a CI check that should fail loudly. Raise it only if your read traffic is bursty; the cap matters because each absorbed failure still costs a backoff, and the overall timeout (10 minutes for waitForRun) keeps ticking.

If the loop does give up, the run is not canceled. Store the run id before you wait, then read it back later or receive the result by run webhook instead. The related guide on SumeRunTimeoutError covers resuming by id.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume