Poll a Sume bulk queue every 10 seconds: read limits by plan
How often can you poll a Sume bulk queue? Reads are 40 times the write limit, so a 10-second poll fits every plan. The math, and when to back off.

Yes, a 10-second poll on a bulk queue is comfortably inside every plan's read limit. Sume rate limits writes and reads per minute, and the read allowance is 40 times the write allowance: Free 120 writes and 4,800 reads, Pro 300 and 12,000, Startup 600 and 24,000, Scale 1,200 and 48,000 (read 2026-10-06, Sume docs: Errors and spend). One queue polled every 10 seconds is six reads a minute.
The number that can bite is not reads but the way you poll. A loop with no sleep, or one that polls every child run instead of the queue, multiplies quickly. Poll the queue once, and open children only when counts say something needs attention.
The arithmetic
Take a season of 100 items. Polling the queue every 10 seconds costs 6 reads a minute regardless of how many items are inside, because the queue endpoint returns counts for all of them. Polling each of the 100 children at the same interval would cost 600 reads a minute, which is 5 percent of a Pro plan's 12,000 but 12.5 percent of a Free plan's 4,800, and it tells you nothing the queue count did not.
The writes side is the one to watch for a bulk job. Creating the queue is one write. Cancelling children is one write each. Retrying 20 failures by single runs would be 20 writes, while one new bulk queue is one.
A polite polling loop
Use the queue as the progress signal and back off as the job ages. Start at 10 seconds, move to 30 after the first few minutes and read children only after the status is completed. Stop on completed, but then read counts.failed, since completed does not mean every item succeeded.
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
URL = "https://api.sume.com/v1/format-run-queues/" + os.environ["QUEUE_ID"]
started = time.time()
while True:
r = requests.get(URL, headers=H, timeout=30)
if r.status_code == 429:
time.sleep(int(r.headers.get("Retry-After", "30")))
continue
q = r.json()["data"]
print(q["status"], q["counts"])
if q["status"] == "completed":
break
time.sleep(10 if time.time() - started < 300 else 30)
print("failed:", q["counts"]["failed"])| Plan | Read limit per minute | One queue at 10 s | 20 queues at 10 s |
|---|---|---|---|
| Free | 4,800 | 6 | 120 |
| Pro | 12,000 | 6 | 120 |
| Startup | 24,000 | 6 | 120 |
| Scale | 48,000 | 6 | 120 |
Prefer a push for long jobs
A queue has no webhook of its own, but each item can carry a webhook_url. For a job that runs for hours, count item events and poll the queue only as a slow backstop, say once a minute. That keeps your reads low and your latency to the finish line short. The same pattern holds for a single run, where the receipt's status_url is the small payload to poll, and result_url is for the full receipt once the run is terminal.
If you do hit a 429, treat it as a signal to slow down, not to retry in a tighter loop. Sleep for the interval the server gives, then resume at the slower cadence. The retry-after header is sent on a 429, and error.details.scope says whether the read or the write budget ran out; the ratelimit-remaining and ratelimit-reset headers let you pace without counting requests yourself.
Writes deserve the same care
Reads are generous, but writes are the scarce side: Free allows 120 a minute and Scale 1,200. A bulk queue turns up to 100 writes into one, which is the main reason to use it. If you find yourself cancelling many children, that is a loop of single writes, so pace it rather than firing all of them at once.
Mind the whole account
Each API key gets its own per-minute budget, sized by the plan of the workspace that owns the key. A dashboard, a monitoring job and your bulk poll that share one key therefore draw from the same pool of reads, while a separate key for each tool keeps them apart. If several tools poll the same queue, consolidate them behind one cache of your own that polls once and serves the rest.
Finally, log the 429s you see. A steady trickle of them means a loop is too tight somewhere, and that is cheaper to find from a log line than from an incident.
Sources
Related posts
More in Formats
- Retry only the failed episodes from a Sume bulk queue
A Sume bulk queue shows completed even when items failed. Find the failed children and re-queue only those, with fresh keys and untouched keys for the rest.
- Season 2 of a Shorts series: same Sume Format, new input, new keys
Start season 2 without rebuilding anything. Keep the Format, change the input and idempotency keys, and keep season 1 receipts as a style reference.
- Shorts series bulk queue finishes out of order: publish by number
Episodes in a Sume bulk queue run in parallel and finish in any order. Sort by an episode number in your output schema, not by completion time.
- Shorts series: episode video and thumbnail from one Format run
YouTube Shorts series take a custom thumbnail per episode. Bind one output schema so each Sume Format run returns the video and the thumbnail together.
Written by Sume