Status reads for a full Sume queue: 3,600 to 72,000

Poll every accepted job every 2 s for 20 minutes: Free makes 3,600 reads, Scale 72,000. The math, and how one field cuts it.

5 min readSume
All posts

The worst case for a naive poll loop is every accepted job polled at the 2-second floor for the whole 20-minute default wait of waitForJob. That is 20 x 60 / 2 = 600 reads per job. Multiply by the accepted job capacity of each plan and you get 3,600 reads for Free (6 jobs), 14,400 for Pro (24), 28,800 for Startup (48) and 72,000 for Scale (120).

These are ceilings, not typical numbers, and they ignore that jobs finish. They are still the right size to design against, because read endpoints can have their own rate limits. The Generation admission page says to treat those as poll backpressure, not as generation concurrency.

Maximum status reads at a 2 s poll over 20 minutes (600 reads per job), as of 2026-10-09
PlanAccepted jobsReads at 2 sReads at 30 s
Free63,600240
Pro2414,400960
Startup4828,8001,920
Scale12072,0004,800

The field that cuts it

The status payload can carry next_poll_after_seconds. The SDK treats its 2-second pollInterval as a floor: when the payload asks for a longer wait, that value wins. When the field is absent, the docs say to use exponential backoff. At 30 seconds the same 20 minutes is 40 reads per job, so a Pro workspace with 24 jobs makes at most 960.

Polling every 2 seconds is rarely useful for video. A job that runs for minutes gains nothing from a second-by-second check, and the cost shows up in your own worker count before it shows up anywhere else.

  • Honor next_poll_after_seconds when it is present.
  • Back off exponentially when it is not, and cap the delay.
  • Stop on terminal: true, not on a status string you guessed.
  • Batch where you can: on MCP, one jobs_wait takes up to 20 job_ids.

Webhook plus a slow poll

A cheaper design is mode: "webhook" with a slow safety poll. The job webhook delivers job.completed, job.failed or job.canceled, and your poll only has to catch what never arrives. Keep it at a minute or more. A webhook is a delivery optimization, and the docs say to keep the status_url poll available for missed deliveries.

For the arithmetic of the other direction, see the 12-minute job example.

Putting a budget in code

Compute a per-worker limit, not a global one. If one process polls 24 jobs, an interval of 30 seconds gives it 24 x 2 = 48 reads a minute, and a 2-second interval gives 720. Choose the interval from the longest job you expect: for a 10-minute video, checking every 30 seconds costs 20 reads per job and rarely delays your reaction by more than half a minute.

Add jitter so many workers do not poll in the same second, and add the retry-after header to your sleep when a read returns 429. If you see rate limits on reads, lengthen the interval before you shorten anything else.

Treat this as a habit, not a one-time fix. Write the rule down next to the code that calls the API, add a test that exercises it, and review it whenever the docs change. Check the linked documentation pages in the sources list for the current wording before you rely on any number here, because limits and field names can be revised, and a short test run costs far less than debugging a production incident.

When something does not match what you read here, capture the x-sume-request-id response header and the job or run id, and send those to support. Do not paste API keys, signing secrets or full request bodies into a ticket or a chat; the ids are enough for the team to find the request.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume