Poll hundreds of AI jobs without a thundering herd: jitter and budgets
Poll many Sume jobs without synchronized bursts: jitter, next_poll_after_seconds, per-plan read budgets, and the math on how much polling a plan can absorb.

Add random jitter to every poll and honor next_poll_after_seconds when the status payload carries it. A thundering herd happens when many workers start together, sleep for the same fixed interval, and then all read at the same instant. Sume gives reads their own rate-limit bucket, so the herd does not block your submits, but it can still earn a 429 on the read side and waste capacity you want for other callers on the same key.
This page shows how much polling each plan absorbs, why a single workspace rarely gets near the ceiling, and a poller that spreads its reads.
What the read budget is
Each API key has a per-minute budget for all of /v1, set by the plan of the workspace that owns the key. Reads (any GET or HEAD, including status_url, events_url and result_url) and writes (submits, cancels, uploads) are counted in separate buckets. The read bucket is forty times the write bucket by default, because agents poll a lot and polling costs nothing.
A 429 names the bucket that ran out in error.details.scope (read or write), and carries retry-after. A 429 on a read means only that the read failed. The job keeps running and keeps billing, so back off and read again rather than treating the job as failed.
How much polling a plan can absorb
The accepted-job ceiling is the number of paid generation jobs a workspace can have processing or queued at once. That ceiling bounds how many jobs you can have in flight, and therefore how many you ever need to poll. The table assumes the worst case: every accepted job polled every 2 seconds, which is 30 reads per minute per job.
| Plan | Reads per minute | Accepted jobs | Reads/min if all polled every 2 s | Share of read budget |
|---|---|---|---|---|
| Free | 4,800 | 6 | 180 | 3.75% |
| Pro | 12,000 | 24 | 720 | 6% |
| Startup | 24,000 | 48 | 1,440 | 6% |
| Scale | 48,000 | 120 | 3,600 | 7.5% |
So where does a herd come from
One workspace rarely exhausts its read bucket with job polls alone, so the herd usually comes from somewhere else. Several services share one key. A deploy restarts fifty workers at once and they all poll on the same beat. A dashboard polls every open tab. A retry loop with no backoff turns one slow response into a burst. Each of those multiplies reads without adding any information, because a job changes state a handful of times in its life.
Two documented behaviors already push the right way. The status payload can carry next_poll_after_seconds, and when it is present you should obey it. The TypeScript helper waitForJob treats the 2-second pollInterval as a floor and lets a longer server hint win, and the SDK helpers add jitter to their polls for the reason described above.
A poller that spreads its reads
The Python below starts each watcher after a random 0 to 2 second delay, sleeps for the server hint when there is one, multiplies the hint by a random factor between 0.8 and 1.2, and keeps a client-side deadline. The deadline is yours: a timeout stops the wait, it does not cancel the job, so store the job id before you raise. It uses only the standard library and the flat status shape (terminal, sume_status, next_poll_after_seconds). Set JOB_IDS to a comma-separated list.
import asyncio, json, os, random, urllib.request
API = "https://api.sume.com"
def read(path: str) -> dict:
req = urllib.request.Request(API + path, headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]})
with urllib.request.urlopen(req, timeout=30) as resp:
return json.load(resp)
async def watch(job_id: str, deadline: float = 1200.0) -> dict:
loop = asyncio.get_running_loop()
end, wait = loop.time() + deadline, 2.0
await asyncio.sleep(random.uniform(0, 2)) # spread the first read
while loop.time() < end:
status = await asyncio.to_thread(read, f"/v1/jobs/{job_id}/status")
if status["terminal"]:
return status
hint = status.get("next_poll_after_seconds") or wait
wait = min(wait * 1.5, 30.0)
await asyncio.sleep(hint * random.uniform(0.8, 1.2))
raise TimeoutError(job_id) # the job still runs and bills
async def main() -> None:
ids = os.environ["JOB_IDS"].split(",")
done = await asyncio.gather(*(watch(i) for i in ids))
print([d["sume_status"] for d in done])
asyncio.run(main())Rules that keep it cheap
- Stop on
terminal, not on a guess. The booleansterminalandresult_readyare the stop signal. - Prefer a webhook for long work, with one slow poll as the backup. A video job that takes minutes does not need a read every two seconds.
- Poll
status, not the full job.GET /v1/jobs/{id}/statusis the lightweight read meant for polling; fetchresultonce, whenresult_readyis true. - On a read
429, wait forretry-afterand keep the job id. Do not resubmit the paid request. - Never put the poll loop inside the request path of your own users. Poll in a worker and let your UI read your own database.
Sources
Related posts
More in Developers
- Portuguese speech to text API: Sume STT language_code pt or pt-BR
Transcribe Portuguese audio with Sume STT using language_code pt or pt-BR, then check the reported language and word times. $0.01 per audio minute.
- Probe a finished video before upload: duration, size and aspect
Run video inspect with frames false to read a render's duration, size and frame rate before posting. Check it against the 3-minute Shorts limit.
- Python: cheapest Sume image model that lists your aspect ratio
A 25-line Python script reads Sume's image catalog, keeps models that list your aspect ratio, prices each from its endpoints record and prints the cheapest.
- Reconcile Sume jobs after a deploy or outage: poll what is open
After downtime, read status for every job your own table still shows as open, honor terminal and result_ready, and never resubmit. Python with sqlite.
Written by Sume