Status reads for a full Sume queue: 3,600 to 72,000
Poll every accepted job every 2 s for 20 minutes: Free makes 3,600 reads, Scale 72,000. The math, and how one field cuts it.

The worst case for a naive poll loop is every accepted job polled at the 2-second floor for the whole 20-minute default wait of waitForJob. That is 20 x 60 / 2 = 600 reads per job. Multiply by the accepted job capacity of each plan and you get 3,600 reads for Free (6 jobs), 14,400 for Pro (24), 28,800 for Startup (48) and 72,000 for Scale (120).
These are ceilings, not typical numbers, and they ignore that jobs finish. They are still the right size to design against, because read endpoints can have their own rate limits. The Generation admission page says to treat those as poll backpressure, not as generation concurrency.
| Plan | Accepted jobs | Reads at 2 s | Reads at 30 s |
|---|---|---|---|
| Free | 6 | 3,600 | 240 |
| Pro | 24 | 14,400 | 960 |
| Startup | 48 | 28,800 | 1,920 |
| Scale | 120 | 72,000 | 4,800 |
The field that cuts it
The status payload can carry next_poll_after_seconds. The SDK treats its 2-second pollInterval as a floor: when the payload asks for a longer wait, that value wins. When the field is absent, the docs say to use exponential backoff. At 30 seconds the same 20 minutes is 40 reads per job, so a Pro workspace with 24 jobs makes at most 960.
Polling every 2 seconds is rarely useful for video. A job that runs for minutes gains nothing from a second-by-second check, and the cost shows up in your own worker count before it shows up anywhere else.
- Honor
next_poll_after_secondswhen it is present. - Back off exponentially when it is not, and cap the delay.
- Stop on
terminal: true, not on a status string you guessed. - Batch where you can: on MCP, one
jobs_waittakes up to 20job_ids.
Webhook plus a slow poll
A cheaper design is mode: "webhook" with a slow safety poll. The job webhook delivers job.completed, job.failed or job.canceled, and your poll only has to catch what never arrives. Keep it at a minute or more. A webhook is a delivery optimization, and the docs say to keep the status_url poll available for missed deliveries.
For the arithmetic of the other direction, see the 12-minute job example.
Putting a budget in code
Compute a per-worker limit, not a global one. If one process polls 24 jobs, an interval of 30 seconds gives it 24 x 2 = 48 reads a minute, and a 2-second interval gives 720. Choose the interval from the longest job you expect: for a 10-minute video, checking every 30 seconds costs 20 reads per job and rarely delays your reaction by more than half a minute.
Add jitter so many workers do not poll in the same second, and add the retry-after header to your sleep when a read returns 429. If you see rate limits on reads, lengthen the interval before you shorten anything else.
Treat this as a habit, not a one-time fix. Write the rule down next to the code that calls the API, add a test that exercises it, and review it whenever the docs change. Check the linked documentation pages in the sources list for the current wording before you rely on any number here, because limits and field names can be revised, and a short test run costs far less than debugging a production incident.
When something does not match what you read here, capture the x-sume-request-id response header and the job or run id, and send those to support. Do not paste API keys, signing secrets or full request bodies into a ticket or a chat; the ids are enough for the team to find the request.
Sources
Related posts
More in Developers
- Polling a 5-minute video: 30 s is 10 reads, 2 s is 150 on Sume
How often to poll a Sume video job: read counts for a 300 s job at 30 s and 2 s, and an asyncio loop that obeys next_poll_after_seconds.
- Porting chat history to Agent Completions: assistant turns fail
Sume Agent Completions take OpenAI-style messages but return 400 on any assistant turn and start each call in a new thread. What to send instead.
- OpenRouter video payload on Sume: drop seed, size, provider.options
Sume /v1/videos follows the OpenRouter shape but rejects seed, size and provider.options with a 400. A Node function that strips them and checks the model id.
- POST /v1/videos status codes: which of 10 are safe to retry
OpenAPI lists 202 plus ten error codes for POST /v1/videos. Which to retry with the same key, which to fix or stop on, a Python classifier and a backoff plan.
Written by Sume