Poll 40 Sume jobs without bursts: jitter, retry-after, next poll hint

Watch many Sume jobs in asyncio: stagger the first poll, sleep next_poll_after_seconds with jitter, and back off on a 429 retry-after from the status route.

5 min readSume
All posts

Start each watcher after a random delay of a few seconds, sleep for the job's next_poll_after_seconds multiplied by a jitter factor between 1.0 and 1.25, and treat a 429 from the status route as a request to slow down, honoring retry-after. Forty jobs submitted in one burst then spread their polls instead of calling the status route in lockstep.

Generation admission says that read, status and list endpoints can have their own rate limits, and that you should treat them as poll backpressure, not generation concurrency. A 429 here means "poll less". It does not mean the job failed, and it never justifies a new submit.

The watcher

The script reads job ids from JOB_IDS, one watcher per id. It stops each watcher when terminal is true, and prints the final sume_status. It needs httpx.

import asyncio, os, random
import httpx

BASE = "https://api.sume.com/v1"

async def watch(client: httpx.AsyncClient, job_id: str) -> tuple[str, str]:
    await asyncio.sleep(random.uniform(0, 5))
    while True:
        r = await client.get(f"{BASE}/jobs/{job_id}/status")
        if r.status_code == 429:
            await asyncio.sleep(float(r.headers.get("retry-after", 5)) + random.random())
            continue
        r.raise_for_status()
        body = r.json()
        s = body.get("data", body)
        if s.get("terminal"):
            return job_id, str(s.get("sume_status"))
        hint = s.get("next_poll_after_seconds") or 10
        await asyncio.sleep(hint * random.uniform(1.0, 1.25))

async def main() -> None:
    key = os.environ.get("SUME_API_KEY", "")
    ids = [j for j in os.environ.get("JOB_IDS", "").split(",") if j]
    if not key or not ids:
        raise SystemExit("set SUME_API_KEY and JOB_IDS=job_a,job_b")
    async with httpx.AsyncClient(headers={"Authorization": f"Bearer {key}"}, timeout=30) as client:
        for job_id, status in await asyncio.gather(*(watch(client, j) for j in ids)):
            print(job_id, status)

asyncio.run(main())

What each delay is for

Each timing choice removes one source of synchronized traffic.

Polling delays for a batch of Sume jobs (Sume docs, read 2026-10-04)
DelayValueReason
Initial stagger0 to 5 seconds, randomWatchers do not all start on the same second
Normal sleepnext_poll_after_seconds times 1.0 to 1.25Obeys the hint and spreads jobs that share it
Hint missing10 secondsA fallback; the docs say to back off when the hint is absent
429 on statusretry-after plus up to 1 secondPoll backpressure, not a failed job
TerminalStopterminal is true for completed, failed and canceled

Cut polling instead of tuning it

For a large batch, cut polling instead of tuning it. Submit with mode: "webhook" and a public HTTPS webhook_url, and keep a slow poll as the backup: the jobs docs recommend continuing to poll as a backup for missed or retried deliveries. A webhook-first batch makes a handful of status reads instead of thousands.

When a watcher sees a failed or canceled job, it is done. The error is on the job record, and a retry means a new request under a new idempotency key, which you decide separately.

Caveats

  • A client timeout does not cancel a job. A stopped watcher leaves the job running and billing, so keep the ids.
  • Do not poll faster than the hint to get results sooner. The job does not finish earlier, and you burn read capacity.
  • The status route returns the booleans terminal and result_ready. Fetch the result when result_ready is true.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume