waitForRun 429 and 503: the streak resets only on a clean read
Sume SDK waitForRun counts consecutive failed reads, resets only on a clean one, and retries the final result read so a late 429 cannot lose it.

In the Sume TypeScript SDK, waitForRun tolerates up to six consecutive failed status reads by default, and only a clean read resets the count. After the run turns terminal it also retries the final result read on transient errors, so a 429 at the very end does not throw away a finished, paid-for result.
The summary is in Waiting for runs and jobs. This page covers the details that decide how it behaves under a flaky network or a rate limit.
What counts as a transient failure?
A failed read is SumeRunRequestError. If the error is retryable (the error envelope carries a retryable flag, and the client layer retries 408, 429, 5xx and transport failures), the helper counts it and backs off. If it is not retryable, or the streak is already at maxTransientFailures, the helper throws.
The docs give the reason: a failed read is not a failed run. The run is still executing and still spending, and throwing would lose the handle to it. Back-off honours retryAfterSeconds from the error when the server sent one.
Why does only a clean read reset the streak?
The code comment explains the choice: alternating failures must not run forever on a budget that never resets. Counting failures that happen to be separated by a success would let a flapping endpoint keep the loop alive indefinitely. Counting them only while they are consecutive means six bad reads in a row end the wait, while a single good read gives you a fresh budget.
Each absorbed failure calls onTransientError(error, count), where count is the current streak. Log it; it is the only signal that the wait is limping.
| Situation | Behavior |
|---|---|
| Retryable read error, streak below 6 | Back off, call onTransientError, read again |
| Retryable read error, streak at 6 | Throw SumeRunRequestError |
| Non-retryable read error (for example 401 or 404) | Throw immediately |
| Clean read | Streak resets to zero |
| Timeline read fails (Format runs) | Reported through onTransientError; the wait continues |
| Run is terminal, result read gets 429 | Retry with back-off up to the same limit |
What happens after the run is terminal?
The final read is pure retrieval. The source comment says a 429 there would otherwise throw away a finished result the caller already paid for, so that read gets its own retry loop with the same maxTransientFailures limit and the same back-off.
Non-retryable failures, such as a 404 for a run id that belongs to another key, still throw. The Formats docs note a run you cannot see reads as one that does not exist.
How do I observe it?
Pass onTransientError and onStatus. The sketch logs each absorbed failure with its request id, which is the value to quote if you need to contact Sume.
import { createSumeClient, waitForRun } from "@sume-com/sdk";
const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });
export function watch(runId: string) {
return waitForRun(runId, {
client,
family: "format",
maxTransientFailures: 6,
onTransientError: (error, count) => {
console.warn("transient read failure", count, error.status, error.requestId);
},
onStatus: (status) => console.log("status", status),
});
}What should I still do myself?
The helper protects reads, not creates. A POST is only retried by the client when it carries an Idempotency-Key, since a replay without one would start and bill a second run. Keep a stable key per intent, and treat a thrown SumeRunRequestError after the streak as a reason to read the run later, not to resubmit.
Sources
Related posts
More in Developers
- Test a webhook endpoint before go-live: a Sume CI gate (Python)
Use POST /v1/webhooks/test-deliveries to fire a signed webhook.test at your deployed URL and fail the deploy unless it answers 2xx. Python script included.
- Four avatar clips a week: Python batch, one idempotency key each
HeyGen's survey ties avatars to consistent posting. Submit four Sume avatar clips a week from one handle with week-stamped keys and queue_full handling.
- What to measure for Sume API jobs: metrics, labels and alerts
A metrics plan for code that calls the Sume API: submit outcome, queue wait, time to terminal, error code and webhook gap, with low-cardinality labels.
- Which Sume API errors should page an engineer: route by category
Route Sume API failures by category: fix-the-input errors go to the caller, quota to finance, queue to a retry, and only internal or unexpected 5xx to on-call.
Written by Sume