12 Wan 3.0 clips in parallel in Python: ThreadPoolExecutor, width 4
A Python batch for Sume: ThreadPoolExecutor at width 4 (Pro concurrency), one Idempotency-Key per item, polling by next_poll_after_seconds. Cost included.

A thread pool sized to your plan's concurrency is the simplest parallel batch against Sume. Twelve 10-second Wan 3.0 clips at 720p cost $15.00, and on Pro (4 processing seats, 24 accepted jobs) all twelve fit, with four running and eight queued at any moment.
Width 4 here is client-side pacing. The docs recommend keeping new in-flight work within the processing cap, even though the API would also accept queued work, so that your own poll traffic stays small and a stall is easy to see.
The batch
Each item gets a stable key, so re-running the script after a crash resubmits nothing twice.
import os, time, requests
from concurrent.futures import ThreadPoolExecutor
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(item):
i, prompt = item
r = requests.post(API + "/v1/videos", timeout=30,
headers={**H, "Idempotency-Key": "launch-7-clip-%03d" % i},
json={"model": "wan-3.0", "prompt": prompt, "duration": 10,
"resolution": "720p", "aspect_ratio": "9:16"})
r.raise_for_status()
job_id = r.json()["id"]
while True:
s = requests.get(API + "/v1/jobs/%s/status" % job_id, headers=H, timeout=30).json()
s = s.get("data", s)
if s["terminal"]:
return i, job_id, s["sume_status"]
time.sleep(s.get("next_poll_after_seconds") or 5)
prompts = ["Product turntable on white, scene %d" % n for n in range(12)]
with ThreadPoolExecutor(max_workers=4) as pool:
for i, job_id, status in pool.map(run, enumerate(prompts)):
print(i, job_id, status)What the script leaves out on purpose
- No retry for 429 or 5xx; wrap the POST in the handler from the retry post when you need it.
- A
402raises fromraise_for_status, which stops that item; catch it per item if you want the rest to continue. pool.mapyields in order, so a slow item 0 delays printing but not the other jobs.
Sizing for other plans
Change max_workers to the effective concurrency_limit from your workspace, not to the table value, because admin overrides and org floors change it. Sixty clips on Pro would exceed the 24 accepted-job capacity, so submit those in waves.
Making it resumable
Because each item has a stable key, you can run the same script again after a crash. Items that were already accepted return their original job; items that were never submitted go through for the first time. That makes the script safe to restart without checking what it did last time, which is the property that matters on a long batch.
- Write each
(index, job_id, status)row to a file as soon as it is known, not only at the end. - On restart, skip rows that already hold a terminal status and let the key return the rest.
- Catch
HTTPErrorper item, record the code, and continue; stop the whole run only on a 402.
If you need async instead of threads, the same shape works with asyncio and a semaphore of the same width. Threads are enough here because each worker spends nearly all its time sleeping between polls.
Sources
Related posts
More in Developers
- Vercel 800 s max duration: do you still need a Sume webhook?
Vercel Pro allows 800 s functions and a 30-minute beta. A Sume video job can still outlast one request, so use async or webhook mode and return in seconds.
- Verify x-sume-webhook-signature in Node: sume-v1 HMAC, raw body
A node:crypto verifier for Sume's sume-v1 signature that refuses an empty secret, checks the 5-minute window, and accepts either entry during a rotation.
- Video upscale hold: omit duration_seconds and Sume reserves 5 s
Sume video upscale reserves from duration_seconds, or 5 seconds when you omit it. Hold table for 5, 15 and 30 seconds at $0.009 per second and the 402 case.
- Wan 3.0 draft and final need different Idempotency-Keys (409)
Reusing one Idempotency-Key for a 480p draft and a 1080p final of the same prompt returns 409 idempotency_conflict. A Python key builder that avoids it.
Written by Sume