12 Wan 3.0 clips in parallel in Python: ThreadPoolExecutor, width 4

A Python batch for Sume: ThreadPoolExecutor at width 4 (Pro concurrency), one Idempotency-Key per item, polling by next_poll_after_seconds. Cost included.

3 min readSume
All posts

A thread pool sized to your plan's concurrency is the simplest parallel batch against Sume. Twelve 10-second Wan 3.0 clips at 720p cost $15.00, and on Pro (4 processing seats, 24 accepted jobs) all twelve fit, with four running and eight queued at any moment.

Width 4 here is client-side pacing. The docs recommend keeping new in-flight work within the processing cap, even though the API would also accept queued work, so that your own poll traffic stays small and a stall is easy to see.

The batch

Each item gets a stable key, so re-running the script after a crash resubmits nothing twice.

import os, time, requests
from concurrent.futures import ThreadPoolExecutor

API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(item):
    i, prompt = item
    r = requests.post(API + "/v1/videos", timeout=30,
                      headers={**H, "Idempotency-Key": "launch-7-clip-%03d" % i},
                      json={"model": "wan-3.0", "prompt": prompt, "duration": 10,
                            "resolution": "720p", "aspect_ratio": "9:16"})
    r.raise_for_status()
    job_id = r.json()["id"]
    while True:
        s = requests.get(API + "/v1/jobs/%s/status" % job_id, headers=H, timeout=30).json()
        s = s.get("data", s)
        if s["terminal"]:
            return i, job_id, s["sume_status"]
        time.sleep(s.get("next_poll_after_seconds") or 5)

prompts = ["Product turntable on white, scene %d" % n for n in range(12)]
with ThreadPoolExecutor(max_workers=4) as pool:
    for i, job_id, status in pool.map(run, enumerate(prompts)):
        print(i, job_id, status)

What the script leaves out on purpose

  • No retry for 429 or 5xx; wrap the POST in the handler from the retry post when you need it.
  • A 402 raises from raise_for_status, which stops that item; catch it per item if you want the rest to continue.
  • pool.map yields in order, so a slow item 0 delays printing but not the other jobs.

Sizing for other plans

Change max_workers to the effective concurrency_limit from your workspace, not to the table value, because admin overrides and org floors change it. Sixty clips on Pro would exceed the 24 accepted-job capacity, so submit those in waves.

Making it resumable

Because each item has a stable key, you can run the same script again after a crash. Items that were already accepted return their original job; items that were never submitted go through for the first time. That makes the script safe to restart without checking what it did last time, which is the property that matters on a long batch.

  • Write each (index, job_id, status) row to a file as soon as it is known, not only at the end.
  • On restart, skip rows that already hold a terminal status and let the key return the rest.
  • Catch HTTPError per item, record the code, and continue; stop the whole run only on a 402.

If you need async instead of threads, the same shape works with asyncio and a semaphore of the same width. Threads are enough here because each worker spends nearly all its time sleeping between polls.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume