waitForJob polls every 2 s: 600 reads in a 20-minute Sume video job
The SDK's waitForJob floor is a 2-second poll and a 20-minute timeout: at most 600 reads per job. The read budget by plan and why next_poll_after_seconds wins.

The Sume TypeScript SDK's waitForJob defaults to a 20-minute timeout and a 2-second poll interval. If the server never asks for a longer wait, that is at most 600 status reads for one job, or 30 per minute. The poll interval is a floor: when the status payload carries next_poll_after_seconds, that value wins, so a well-behaved job costs fewer reads.
Reads have their own budget, forty times the write number, so a poll loop does not block your submits. The arithmetic below shows how many jobs you could poll at the floor before you reach the read limit.
Budget math
Per-minute read budgets come from the authentication docs. Jobs-at-floor is the budget divided by 30 reads a minute per job, rounded down. Accepted capacity is from the generation admission docs; it is the number of paid jobs that can be queued or processing at once.
| Plan | Reads per minute | Jobs pollable at 2 s | Accepted jobs |
|---|---|---|---|
| Free | 4800 | 160 | 6 |
| Pro | 12000 | 400 | 24 |
| Startup | 24000 | 800 | 48 |
| Scale | 48000 | 1600 | 120 |
What the table says
On every plan, accepted capacity is far smaller than the number of jobs you could poll at the floor. So for generation jobs a plain waitForJob loop per job is safe on reads. The docs' warning is about tight loops across many jobs; honor next_poll_after_seconds and stop on a terminal status.
Formats are different in one respect: timeline: true on subscribeFormatRun or waitForRun also reads the phase timeline on every poll, which doubles the request rate of the wait. It is off by default for that reason.
Using it
The helper resolves with the job record, not the /result body. A failed or canceled job gives 409 job_not_completed on /result, so read status, result and error from the record. A client-side timeout throws SumeJobTimeoutError carrying the jobId; it does not cancel the job, which keeps running and billing, so store the id and resume or cancel.
Submit with async or no mode. The server-side sync and subscribe modes wait for at most 30 seconds, which is a budget for the HTTP request and not for the job.
import { createSumeClient, generateVideoV1, waitForJob } from "@sume-com/sdk";
const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! });
const { data, error } = await generateVideoV1({
client,
headers: { "idempotency-key": crypto.randomUUID() },
body: { prompt: "Slow push-in on a ceramic mug", mode: "async" },
});
if (error) throw new Error(JSON.stringify(error));
const job = await waitForJob(data!.data.request_id, {
client,
timeout: 20 * 60_000,
pollInterval: 2_000,
});
console.log(job.status);
When to use a webhook instead
A webhook removes the poll loop entirely, and the docs recommend keeping polling as a backup. For a few concurrent jobs the SDK helper is the shorter path; for hundreds, register a callback URL, verify it, and let a sweeper poll only jobs that overshoot their expected time. Either way, keep the job id in your database before you start waiting.
Browsers should not poll your Sume key. Poll from your backend and push state to the client.
A budget to write down
Put three numbers in your runbook: the maximum jobs in flight, the poll interval you chose, and the reads per minute those two imply. Compare the third against your plan's read budget with a margin of at least half. If the margin disappears, move the oldest jobs to the sweeper schedule first.
Sources
Related posts
More in Developers
- Wan 3.0 API request cheat sheet: three modes, 2 to 30 seconds
Wan 3.0 on Sume: the request body for text, first/last frame and reference modes, the 480p/720p/1080p rates and the 2 to 30 second window, on one page.
- Wan 3.0 API rate limit: 300 RPM on Model Studio vs Sume jobs
Alibaba lists 300 requests per minute for Wan 3.0 on Model Studio. Here is what that means for a batch, and how Sume submits Wan 3.0 as async jobs.
- Hackathon app on the Sume Free plan: 1 seat, 6 accepted jobs
A weekend demo on Free can have one job processing and five queued. How to design the UI, the retries and the demo script around that, with the real limits.
- GPT Image 1 to GPT Image 2.5 on Sume: what changes in the output
Moving from GPT Image 1 to ChatGPT Image 2.5 on Sume changes the response (URL, not base64), default quality, size grid and failures.
Written by Sume