asyncio TaskGroup cancels siblings: poll many Sume jobs safely
A TaskGroup cancels every other poller when one raises. For a batch of Sume video jobs that abandons waits, not jobs. Catch inside the task and return results.

asyncio.TaskGroup cancels every other task in the group as soon as one task raises, so polling a batch of Sume jobs with a bare TaskGroup means one failed job abandons the waits on all the others. The jobs themselves keep running and billing on Sume's side, since a client-side cancel does not cancel the job, per Sume jobs and results. You lose the handles, not the charges. The fix is to catch errors inside each task and return them as values, so the group only ever sees tasks that finish normally. The group still waits for every task, which is what you want: the call to async with returns only when each job has either reached a terminal state, hit your deadline, or produced a recorded error.
This is easy to miss when porting a sequential loop to concurrent polling, because the happy path works and the first failure arrives in production.
Why is that the wrong default for polling?
Python's documentation describes the behaviour: when a task in a TaskGroup fails with an exception other than CancelledError, the remaining tasks are cancelled and the group raises an ExceptionGroup once it unwinds. For a poller that is the wrong default. A failed job is an ordinary outcome that you want to record next to the successes, with its error category and retryable flag, not an emergency that stops its neighbours.
What does the safe version look like?
import asyncio
async def poll_one(job_id, fetch):
try:
for _ in range(100):
status = await fetch(job_id)
if status["terminal"]:
return job_id, status
await asyncio.sleep(status.get("next_poll_after_seconds", 0.01))
return job_id, {"sume_status": "deadline", "terminal": False}
except Exception as exc:
return job_id, {"sume_status": "poll_error", "error": repr(exc)}
async def main():
script = {"a": ["processing", "completed"], "b": ["failed"], "c": ["processing", "completed"]}
async def fetch(job_id):
state = script[job_id].pop(0)
if job_id == "b":
raise RuntimeError("status endpoint blew up")
return {"sume_status": state, "terminal": state == "completed"}
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(poll_one(j, fetch)) for j in script]
for task in tasks:
print(task.result())
assert sum(1 for t in tasks if t.result()[1]["sume_status"] == "completed") == 2
asyncio.run(main())What happened to job b?
Job b raises inside fetch, and jobs a and c still complete, because poll_one converted the exception into a returned value. Notice the third outcome, deadline: a poller that gives up must say so, and the caller must record the job id, because that job may still run. A poll error is also not a job failure. A transient 5xx on a status read means retry the read, not that the job died.
| Outcome | How the poller reports it | What you do next |
|---|---|---|
| completed | Status with terminal true | Fetch the result |
| failed | Status with terminal true | Read the error, decide on a retry |
| deadline | Not terminal | Store the id, read it later or cancel |
| poll error | Exception turned into a value | Retry the read with backoff |
Do I still need a semaphore?
Bound the concurrency as well. A TaskGroup with 500 tasks sends 500 status reads at once, and a flood of reads is wasteful and can draw rate limits, which are retryable, not job failures. Wrap fetch in an asyncio.Semaphore of 10 or 20, add jitter to the sleep, and honour next_poll_after_seconds when it is present. The wave size for submits is a separate number: the admission fields (see Sume generation admission) queue_capacity_remaining and wave_size_hint tell you how many jobs to start together. Treat those two numbers as hints, not guarantees, you re-read before each wave rather than constants to cache, because capacity changes while your batch runs. Finally, log the job id before the first poll, not after the first result: if the process dies mid-batch, the ids you stored are the only way to find the clips you already paid for, and the same stored ids let a restarted run resume polling instead of resubmitting. Resubmitting without the original idempotency key is exactly how a batch of 40 becomes a bill for 80.
Two smaller habits keep the batch honest. First, give every task its own deadline measured from a monotonic clock, not a count of loops, so a slow job cannot stretch the whole batch; the Sume docs suggest a client-side deadline such as 20 minutes for video, and note that it is separate from wait_timeout_seconds, which never exceeds 30 seconds (read 2026-10-06). Second, when the group finishes, write one line per job id with its final state, including deadline and poll_error, before you do anything else with the results. That file is what lets a later run poll the leftovers instead of paying again.
Sources
Related posts
More in Developers
- Check balance before bulk transcription: 402 and admission preview
A 402 insufficient_credits arrives before any provider work. Read GET /v1/balance and POST /v1/generation/admission-preview first and size the run.
- Port a Bedrock image call to Sume /v1/images in Python
Moving from boto3 invoke_model for Nova Canvas or Titan to Sume's REST call: the request mapping, the response shape, the status codes, and the swap code.
- BullMQ delayed job that polls an AI video job and reschedules itself
A BullMQ worker reads Sume's job status once, then adds the next poll with a delay from next_poll_after_seconds, so no worker slot is held while a clip renders.
- A calendar file for AI model shutdown dates: .ics from Python
Generate an .ics file with all-day events and 14-day reminders for the gpt-image-1 and gpt-image-1.5 shutdown dates, then import it into any calendar app.
Written by Sume