Python async generator that yields Sume job status until terminal
Write an async generator in stdlib Python that polls /v1/jobs/:id/status, honours next_poll_after_seconds, and stops at a terminal status. No SDK needed.

Wrap the status call in an async def that loops, yields each snapshot, and returns once the response says terminal or the status is completed, failed or canceled. Sume publishes no Python SDK, so the transport is plain HTTP. Run the blocking urllib call inside asyncio.to_thread, and the generator stays friendly to other tasks in the same event loop.
Why a generator fits polling
The payoff of a generator is that the caller decides what to do with each step. A web handler can stream progress to a browser, a script can print it, and a test can feed a fake fetch function and assert on the sequence. The poll loop itself never needs to know.
- The status route is
GET https://api.sume.com/v1/jobs/{job_id}/statuswith one credential header, eitherAuthorization: Bearerorx-api-key, never both. - Stop on
completed,failedorcanceled. A job inqueuedis not a failure, because Sume queues work when your workspace is at its concurrency limit. - The response carries
next_poll_after_seconds. Wait at least that long between polls, and fall back to a floor of 2 seconds. - A timeout in your own process is not a failed job. Resume with the same id. Do not submit the paid request again.
- Each snapshot also keeps the queue shaped
statusfield, which mirrorssume_status, so log both when you debug a stuck job.
The generator
The sample takes the fetch function as an argument. The real one is below it, and a fake one makes the file run offline.
import asyncio, json, os, urllib.request
TERMINAL = {"completed", "failed", "canceled"}
def http_status(job_id):
req = urllib.request.Request(
f"https://api.sume.com/v1/jobs/{job_id}/status",
headers={"x-api-key": os.environ["SUME_API_KEY"]})
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)
async def watch(job_id, fetch=http_status, floor=2.0):
while True:
snap = await asyncio.to_thread(fetch, job_id)
yield snap
if snap.get("terminal") or snap.get("status") in TERMINAL:
return
await asyncio.sleep(max(floor, snap.get("next_poll_after_seconds") or 0))
async def demo():
steps = iter([{"status": "queued"}, {"status": "processing"}, {"status": "completed", "terminal": True}])
async for snap in watch("job_demo", fetch=lambda _: next(steps), floor=0.01):
print(snap["status"])
asyncio.run(demo())
After the loop ends
The generator hands control back to the caller on every step, and that is the point. A FastAPI endpoint can forward each snapshot over server sent events. A CLI can print a dot per poll. A test can pass a lambda that returns canned dictionaries, as the demo does, and check that the loop stops at the right time without any network or any sleeping.
After the generator ends, read the last snapshot. completed means you can call GET /v1/jobs/{id}/result. For any other terminal status, the result route answers 409 job_not_completed, so read the failure from the job record instead. A failed job carries an error with a category such as validation, quota, generation_timeout or internal, plus a retryable flag.
Deadlines and cancellation
Call the generator from asyncio.run, as in the demo, or from inside a running loop with async for. There is no top level await in a plain script. If you want a hard deadline, wrap the loop in asyncio.timeout, and treat the timeout as a decision to stop watching, not as a job failure. The SDK helper waitForJob uses a 20 minute default for the same reason, and it throws a timeout error that does not cancel the job.
Rules to add before production
The generator above handles the happy path. Production code needs a few more rules, and each of them comes from how the API behaves, not from taste.
- Handle read rate limits. Status polls are reads, and each key has its own read bucket per minute, separate from the write bucket. A 429 names
error.details.scopeand carriesretry-after, so sleep for that value. - Do not poll in a tight loop for many jobs at once. One generator per job multiplies reads, so share one scheduler when you watch hundreds of ids, or prefer a webhook with a poll as the backup.
- Treat transient HTTP failures as a reason to retry the read. A status read is safe to repeat, and only the paid submit call needs an idempotency key.
- Log the
request_idfrom any error envelope. Support can find a request by that id, and it is safe to paste into a ticket.
Store the result once
If you need the finished artifact URL, fetch the result once. The artifact URL is durable and public on media.sume.com, so store it in your database and stop reading the job. A later read of the same job is not needed to show the asset to a customer.
Sources
Related posts
More in Developers
- Python: price a Seedance 2.5 clip from seconds and resolution
A 17-line function turns resolution and seconds into video tokens and a Sume price: 4 s at 480p is $1.07, 30 s at 1080p is $42.65.
- Python: quote one 10-second clip across five Sume video models
A Python script that prices a 10-second clip on Seedance 2.5, Omni, Wan 3.0, H3 and H3 Max with Sume's 1.25 multiple and per-job rounding. Run it before a test.
- Python requests Retry on POST: safe for Sume only with a key
urllib3 Retry skips POST by default. Allow it for a Sume submit only when every request carries an Idempotency-Key, and retry just 429 and 503.
- Flag near-identical openings in your last 20 Short scripts
A short Python script that compares the first sentence of each Short script and flags pairs that read alike, a cheap check against templated AI Shorts.
Written by Sume