Python ThreadPoolExecutor: submit 12 Wan 3.0 clips, four at a time
Stdlib-only Python: ThreadPoolExecutor with 4 workers posts Wan 3.0 jobs to /v1/videos, one Idempotency-Key per shot, results printed in completion order.

A submit call is blocking network I/O, so threads are a fine fit and you need no async runtime. concurrent.futures.ThreadPoolExecutor caps how many submits are in flight, and as_completed hands each result back the moment it lands. That second part is the one to remember: completion order is not shot order, so every result must carry its own index.
The sample uses only the standard library (urllib), so it runs on a bare Python 3.14.7. It was tested against a local stand-in server that answers 202, not the live API. Set SUME_BASE (https://api.sume.com) and SUME_API_KEY first.
The code (18 lines)
urllib raises HTTPError for any 4xx or 5xx, so the except branch is where a 429 queue_full, 402 insufficient_credits or 503 provider_capacity_exceeded shows up. The Idempotency-Key is derived from the shot number, so rerunning the script replays jobs that were already accepted.
import json, os, urllib.request
from concurrent.futures import ThreadPoolExecutor, as_completed
def submit(i):
body = json.dumps({"model": "wan-3.0", "prompt": f"shot {i}", "duration": 5, "resolution": "480p"})
req = urllib.request.Request(os.environ["SUME_BASE"] + "/v1/videos", body.encode(), {
"x-api-key": os.environ["SUME_API_KEY"], "content-type": "application/json",
"Idempotency-Key": f"trailer-v3-shot-{i:02d}"})
with urllib.request.urlopen(req, timeout=30) as res:
return i, res.status, json.load(res)["id"]
with ThreadPoolExecutor(max_workers=4) as pool: # 4 = Pro processing concurrency
futures = [pool.submit(submit, i) for i in range(12)]
for f in as_completed(futures): # completion order, not shot order
try:
print(f.result())
except Exception as exc: # urllib raises HTTPError on 4xx/5xx
print("failed:", exc)Choosing max_workers
Four matches the processing concurrency of a Pro workspace. Because Sume queues work it cannot start yet, the real ceiling for submits is the accepted capacity (24 on Pro), so max_workers is about pacing and memory, not about avoiding errors. Read the live numbers from the generation_limits object in a submit response when you need them; the docs label wave_size_hint a hint, not a concurrency limit.
Twelve shots at 480p and 5 seconds is a cheap test: Sume's pricing table carries Wan 3.0 at a provider list of $0.05 a second at 480p, and every SKU sells at list x 1.25, so each clip is $0.3125 and the batch comes to $3.75.
Reading the output
Each printed tuple is (shot, http_status, job_id). Store the pair of shot number and job id, then poll GET /v1/videos/{id} for each one. A shot that printed failed: should be resubmitted with the same key, which is safe because a replay returns the original job.
Sources
Related posts
More in Developers
- Python urllib: honor retry-after on a Sume 429, back off without it
A stdlib retry for GET calls that sleeps for retry-after when the 429 carries it and for a capped exponential delay when it does not, with jitter.
- Python urllib: POST /v1/images, 200 or 202, after Imagen 4 Fast
imagen-4.0-fast-generate-001 ended Aug 17. A stdlib Python call to Sume's image route that reads the status code, then polls the job when the answer is 202.
- queue_full 429 on a Sume submit: the reservation is released
A 429 queue_full releases or refunds the failed admission's reservation. Check refunded_usd_micros in /v1/usage, then retry with the same Idempotency-Key.
- Can I submit 100 AI video jobs at once? Queue limits by plan
Accepted capacity is slots plus queue: 6 on Free, 24 on Pro, 48 on Startup, 120 on Scale. Submit 100 at once and 94, 76, 52 or 0 get 429 queue_full.
Written by Sume