Poll hundreds of AI jobs without a thundering herd: jitter and budgets

Poll many Sume jobs without synchronized bursts: jitter, next_poll_after_seconds, per-plan read budgets, and the math on how much polling a plan can absorb.

5 min readSume
All posts

Add random jitter to every poll and honor next_poll_after_seconds when the status payload carries it. A thundering herd happens when many workers start together, sleep for the same fixed interval, and then all read at the same instant. Sume gives reads their own rate-limit bucket, so the herd does not block your submits, but it can still earn a 429 on the read side and waste capacity you want for other callers on the same key.

This page shows how much polling each plan absorbs, why a single workspace rarely gets near the ceiling, and a poller that spreads its reads.

What the read budget is

Each API key has a per-minute budget for all of /v1, set by the plan of the workspace that owns the key. Reads (any GET or HEAD, including status_url, events_url and result_url) and writes (submits, cancels, uploads) are counted in separate buckets. The read bucket is forty times the write bucket by default, because agents poll a lot and polling costs nothing.

A 429 names the bucket that ran out in error.details.scope (read or write), and carries retry-after. A 429 on a read means only that the read failed. The job keeps running and keeps billing, so back off and read again rather than treating the job as failed.

How much polling a plan can absorb

The accepted-job ceiling is the number of paid generation jobs a workspace can have processing or queued at once. That ceiling bounds how many jobs you can have in flight, and therefore how many you ever need to poll. The table assumes the worst case: every accepted job polled every 2 seconds, which is 30 reads per minute per job.

Read budget versus worst-case polling, default plan limits (read 2026-10-07)
PlanReads per minuteAccepted jobsReads/min if all polled every 2 sShare of read budget
Free4,80061803.75%
Pro12,000247206%
Startup24,000481,4406%
Scale48,0001203,6007.5%

So where does a herd come from

One workspace rarely exhausts its read bucket with job polls alone, so the herd usually comes from somewhere else. Several services share one key. A deploy restarts fifty workers at once and they all poll on the same beat. A dashboard polls every open tab. A retry loop with no backoff turns one slow response into a burst. Each of those multiplies reads without adding any information, because a job changes state a handful of times in its life.

Two documented behaviors already push the right way. The status payload can carry next_poll_after_seconds, and when it is present you should obey it. The TypeScript helper waitForJob treats the 2-second pollInterval as a floor and lets a longer server hint win, and the SDK helpers add jitter to their polls for the reason described above.

A poller that spreads its reads

The Python below starts each watcher after a random 0 to 2 second delay, sleeps for the server hint when there is one, multiplies the hint by a random factor between 0.8 and 1.2, and keeps a client-side deadline. The deadline is yours: a timeout stops the wait, it does not cancel the job, so store the job id before you raise. It uses only the standard library and the flat status shape (terminal, sume_status, next_poll_after_seconds). Set JOB_IDS to a comma-separated list.

import asyncio, json, os, random, urllib.request

API = "https://api.sume.com"

def read(path: str) -> dict:
    req = urllib.request.Request(API + path, headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=30) as resp:
        return json.load(resp)

async def watch(job_id: str, deadline: float = 1200.0) -> dict:
    loop = asyncio.get_running_loop()
    end, wait = loop.time() + deadline, 2.0
    await asyncio.sleep(random.uniform(0, 2))  # spread the first read
    while loop.time() < end:
        status = await asyncio.to_thread(read, f"/v1/jobs/{job_id}/status")
        if status["terminal"]:
            return status
        hint = status.get("next_poll_after_seconds") or wait
        wait = min(wait * 1.5, 30.0)
        await asyncio.sleep(hint * random.uniform(0.8, 1.2))
    raise TimeoutError(job_id)  # the job still runs and bills

async def main() -> None:
    ids = os.environ["JOB_IDS"].split(",")
    done = await asyncio.gather(*(watch(i) for i in ids))
    print([d["sume_status"] for d in done])

asyncio.run(main())

Rules that keep it cheap

  • Stop on terminal, not on a guess. The booleans terminal and result_ready are the stop signal.
  • Prefer a webhook for long work, with one slow poll as the backup. A video job that takes minutes does not need a read every two seconds.
  • Poll status, not the full job. GET /v1/jobs/{id}/status is the lightweight read meant for polling; fetch result once, when result_ready is true.
  • On a read 429, wait for retry-after and keep the job id. Do not resubmit the paid request.
  • Never put the poll loop inside the request path of your own users. Poll in a worker and let your UI read your own database.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume