Python asyncio.timeout around a Sume job poll: a hard budget

Wrap a Sume status loop in asyncio.timeout so it stops at a fixed budget and returns still_running, leaving the job alone. A short version, run against a mock.

4 min readSume
All posts

Put the whole poll loop inside async with asyncio.timeout(budget_s) and catch TimeoutError outside it. The budget then covers every sleep and every blocking read together, and a timeout only ends your wait: the Sume job keeps running and you can read it later with the same status URL. Python 3.11 added asyncio.timeout, so this works on any current interpreter from 3.11 up.

Below, the blocking urllib read runs in asyncio.to_thread so the event loop stays free, and the sleep follows the server's next_poll_after_seconds rather than a fixed interval.

Facts used

Checked 2026-10-02
ItemSource / detail
Status routeGET /v1/jobs/{id}/status, data.terminal and data.sume_status
Poll hintdata.next_poll_after_seconds, null when terminal
Terminal statusescompleted, failed, canceled
asyncio.timeoutPython 3.11 and later

The poller

STATUS_URL is the status_url from the submit response; SUME_API_KEY holds your key.

import asyncio, json, os, urllib.request

def _get(url):
    req = urllib.request.Request(url, headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]})
    with urllib.request.urlopen(req, timeout=20) as r:
        return json.load(r)["data"]

async def wait_job(status_url, budget_s=300):
    try:
        async with asyncio.timeout(budget_s):
            while True:
                d = await asyncio.to_thread(_get, status_url)
                if d["terminal"]:
                    return d["sume_status"]
                await asyncio.sleep(d.get("next_poll_after_seconds") or 2)
    except TimeoutError:
        return "still_running"

async def main():
    print(await wait_job(os.environ["STATUS_URL"], 30))

asyncio.run(main())

Behaviour I observed

Cancelling an in-flight to_thread call does not interrupt the thread: the current read finishes in the background (here capped by its 20 second timeout), then its result is discarded.

  • Against a local mock that completes on the second poll it printed completed.
  • Against one that never completes, with a 1 second budget, it printed still_running after about 1.0 seconds.
  • Run on Python 3.14.7. The mock is my own, not the Sume API.

Choosing the budget

Pick the budget from the caller, not from the job. A request handler that must answer in 25 seconds gets a 20 second budget and returns the status URL when it gets still_running; a background worker can wait several minutes. Whatever you choose, store the job id before the first poll so a restart can resume from the status URL instead of resubmitting and paying twice.

If you submit many jobs, create one wait_job task per job inside a TaskGroup and give each its own budget; one slow job then cannot starve the others.

Limits

A timeout is not a cancel. If you want the job stopped, call the cancel route, which only works before generation starts. The loop treats any HTTP or network error as fatal and will raise; add a retry for 429 and 5xx before using it in production. The budget includes retries you add, so size it to the model's usual duration, not to the poll interval.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume