Poll a Sume bulk queue every 10 seconds: read limits by plan

How often can you poll a Sume bulk queue? Reads are 40 times the write limit, so a 10-second poll fits every plan. The math, and when to back off.

4 min readSume
All posts

Yes, a 10-second poll on a bulk queue is comfortably inside every plan's read limit. Sume rate limits writes and reads per minute, and the read allowance is 40 times the write allowance: Free 120 writes and 4,800 reads, Pro 300 and 12,000, Startup 600 and 24,000, Scale 1,200 and 48,000 (read 2026-10-06, Sume docs: Errors and spend). One queue polled every 10 seconds is six reads a minute.

The number that can bite is not reads but the way you poll. A loop with no sleep, or one that polls every child run instead of the queue, multiplies quickly. Poll the queue once, and open children only when counts say something needs attention.

The arithmetic

Take a season of 100 items. Polling the queue every 10 seconds costs 6 reads a minute regardless of how many items are inside, because the queue endpoint returns counts for all of them. Polling each of the 100 children at the same interval would cost 600 reads a minute, which is 5 percent of a Pro plan's 12,000 but 12.5 percent of a Free plan's 4,800, and it tells you nothing the queue count did not.

The writes side is the one to watch for a bulk job. Creating the queue is one write. Cancelling children is one write each. Retrying 20 failures by single runs would be 20 writes, while one new bulk queue is one.

A polite polling loop

Use the queue as the progress signal and back off as the job ages. Start at 10 seconds, move to 30 after the first few minutes and read children only after the status is completed. Stop on completed, but then read counts.failed, since completed does not mean every item succeeded.

import os, time, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
URL = "https://api.sume.com/v1/format-run-queues/" + os.environ["QUEUE_ID"]

started = time.time()
while True:
    r = requests.get(URL, headers=H, timeout=30)
    if r.status_code == 429:
        time.sleep(int(r.headers.get("Retry-After", "30")))
        continue
    q = r.json()["data"]
    print(q["status"], q["counts"])
    if q["status"] == "completed":
        break
    time.sleep(10 if time.time() - started < 300 else 30)
print("failed:", q["counts"]["failed"])
Reads per minute at each poll interval versus plan read limit, read 2026-10-06 against Sume docs
PlanRead limit per minuteOne queue at 10 s20 queues at 10 s
Free4,8006120
Pro12,0006120
Startup24,0006120
Scale48,0006120

Prefer a push for long jobs

A queue has no webhook of its own, but each item can carry a webhook_url. For a job that runs for hours, count item events and poll the queue only as a slow backstop, say once a minute. That keeps your reads low and your latency to the finish line short. The same pattern holds for a single run, where the receipt's status_url is the small payload to poll, and result_url is for the full receipt once the run is terminal.

If you do hit a 429, treat it as a signal to slow down, not to retry in a tighter loop. Sleep for the interval the server gives, then resume at the slower cadence. The retry-after header is sent on a 429, and error.details.scope says whether the read or the write budget ran out; the ratelimit-remaining and ratelimit-reset headers let you pace without counting requests yourself.

Writes deserve the same care

Reads are generous, but writes are the scarce side: Free allows 120 a minute and Scale 1,200. A bulk queue turns up to 100 writes into one, which is the main reason to use it. If you find yourself cancelling many children, that is a loop of single writes, so pace it rather than firing all of them at once.

Mind the whole account

Each API key gets its own per-minute budget, sized by the plan of the workspace that owns the key. A dashboard, a monitoring job and your bulk poll that share one key therefore draw from the same pool of reads, while a separate key for each tool keeps them apart. If several tools poll the same queue, consolidate them behind one cache of your own that polls once and serves the rest.

Finally, log the 429s you see. A steady trickle of them means a loop is too tight somewhere, and that is cheaper to find from a log line than from an incident.

Sources

Related posts

More in Formats

All Formats posts

Written by Sume