waitForRun and 429 or 503 on a status read: it keeps polling
A 429 or 5xx on a status read does not fail the run. waitForRun tolerates 6 consecutive transient read failures and keeps waiting. Options and what to log.

When a status read returns 429 or a 5xx, the run itself has not failed, and waitForRun does not treat it that way. The SDK tolerates 6 consecutive transient read failures by default, backing off and polling again, because losing the handle to a live run is far more expensive than waiting another second. Raise or lower the tolerance with maxTransientFailures and log each one with onTransientError.
Sources: Waiting for runs and jobs and Errors and spend, read 2026-10-02.
Why is a failed read not a failed run?
The Formats docs are explicit: a 429 or 503 in a poll loop is transient. The read budget is spent, or Sume had a short upstream outage, but the run keeps executing and its spend keeps accruing. Abandoning the loop does not stop it. Bulk queues follow the same rule: a 429 or 503 while polling a queue means the queue is still draining.
| Status | Code | Run state | Do this |
|---|---|---|---|
| 429 | rate_limited (scope: read) | Still running | Wait retry-after, poll again |
| 503 | studio_agent_upstream_unavailable | Still running | Back off, poll again |
| 404 | format_run_not_found | Unknown id or someone else's run | Stop; check the id and the key |
| 403 | insufficient_scope | Not readable with this key | Stop; mint a key with formats:read |
How do I wire the options?
Pass onTransientError so repeated blips show up in your logs instead of vanishing. The docs do not name the error thrown once the limit is passed, so catch any error around the wait. The run is still alive, so record the id and resume from result_url or start a fresh wait.
import { createSumeClient, waitForRun } from "@sume-com/sdk";
const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });
const run = await waitForRun(runId, {
client,
family: "format",
timeout: 15 * 60_000,
pollInterval: 2_000,
maxTransientFailures: 6,
onTransientError: (err) => console.warn("transient read failure", err),
});
console.log(run.status);Why does the poll not stay in lockstep?
Polls are jittered. Reads have their own budget, but several clients started together would otherwise stay in phase and hit that ceiling as a group. The deadline is also checked before sleeping, not after, so a caller who asks for a 5-second timeout hears about it in 5 seconds rather than 5 seconds plus one poll interval.
The failure counter counts consecutive failures, so one good read resets the picture. A long outage that outlasts the tolerance ends the wait, not the run.
What should I do when the wait gives up?
Treat that as a handoff, not a failure. Persist runId, mark it pending, and have a scheduled job resume it. Better still, ask for a run webhook so you are told on completion, and keep this loop as the backup.
Do not start a second paid run for the same intent. If you must retry a create, reuse the same Idempotency-Key so you get the original receipt back.
Sources
Related posts
More in Developers
- Wan 2.2 TI2V-5B in Diffusers: 121 frames, 24 fps, 24GB GPU
A working Diffusers example for Wan2.2-TI2V-5B: 704x1280 portrait frames, 121 frames at 24 fps, 24GB VRAM minimum, and when to use a hosted route instead.
- Wan 3.0 polling every 15 seconds vs Sume's 30-second poll
Alibaba recommends polling Wan 3.0 every 15 seconds. Sume's video guide uses 30 seconds and a signed callback_url. Which to pick and why.
- Wan 3.0 has six regional endpoints; Sume has one base URL
Alibaba's Wan 3.0 API is served from six regions with a workspace id in the URL. Sume's video API uses api.sume.com/v1/videos with no region field.
- Wan 3.0 video URL purged after 24 hours: what Sume returns
Alibaba keeps Wan 3.0 task ids and video URLs for 24 hours. Sume returns media.sume.com artifacts instead. What to store and when to download.
Written by Sume