Can polling hit the Sume read limit? 24 Pro jobs at 2 s use 6%

Polling every accepted job every 2 seconds uses 3.75% to 7.5% of a Sume plan's read budget. The arithmetic for Free, Pro, Startup and Scale, with the caveats.

5 min readSume
All posts

For paid generation jobs, the read limit is not the one you will hit. A Pro workspace can hold 24 accepted jobs (4 processing plus 20 queued). Polling every one of them every 2 seconds is 24 x 30 = 720 reads a minute, which is 6% of the 12,000 reads a minute that Pro gets. Concurrency and queue capacity bind long before the read budget does.

Sume gives each key separate budgets for reads and writes, so a tight poll loop cannot cause a 429 on your submits. Reads are deliberately generous: forty times the write number.

The numbers by plan

Reads per minute come from the Authentication page. Accepted job capacity (concurrency plus default queue capacity) comes from Generation admission. Polling at 2 seconds is 30 polls per job per minute. The last two columns are my arithmetic.

Read budget against polling every accepted job at 2 s, as of 2026-10-09 (Authentication; Generation admission).
PlanReads per minuteAccepted jobsPolls per minute (jobs x 30)Share of read budget
Free4,80061803.75%
Pro12,000247206%
Startup24,000481,4406%
Scale48,0001203,6007.5%

Where the budget can still run out

The table covers paid generation jobs polled at the default cadence. Three things change it. First, other reads count against the same bucket: a poll of status_url, events_url, or result_url, any list call, and the MCP endpoint itself are all reads. Second, Format runs wait differently: with timeline: true, the SDK helpers also read the phase timeline on every poll and double the request rate of the wait. Third, many keys do not add up to one budget, but many processes sharing one key do.

Unauthenticated requests are limited per client IP at the Free rate, with a read bucket of four times the write rate, not forty. Do not poll without a key.

A worked example

Take a Pro workspace that submits 24 image jobs at once. All 24 are accepted: 4 move to processing and 20 sit in queued. A poller that wakes every 2 seconds and reads each job makes 24 reads per cycle, so 720 per minute. Add a dashboard that lists jobs once a second (60 more reads) and a second worker that polls the same 24 jobs (another 720), and the total is 1,500 reads a minute, or 12.5% of 12,000. The write budget is the tighter one: Pro allows 300 writes a minute, and 24 submits are 8% of it.

The practical limit is therefore not request volume. It is that only 4 jobs process at once on Pro and a 25th submit returns 429 queue_full. Size your waves with the generation_limits snapshot instead of by polling cost.

Let the server set the pace

The cheapest way to stay far from the limit is to honor next_poll_after_seconds when it is present, and use exponential backoff with jitter otherwise. The response headers ratelimit-limit, ratelimit-remaining, ratelimit-reset, and retry-after show where you stand, and the docs say not to count requests yourself. The snippet reproduces the table so you can try your own interval.

plans = {  # name: (reads per minute, accepted generation jobs)
    "Free": (4800, 6),
    "Pro": (12000, 24),
    "Startup": (24000, 48),
    "Scale": (48000, 120),
}
interval = 2  # seconds between polls of one job

for name, (reads, jobs) in plans.items():
    polls = jobs * (60 // interval)
    print(f"{name}: {polls} polls/min = {polls / reads:.2%} of {reads}")

Sources

Related posts

More in Developers

All Developers posts

Written by Sume