Python: run one prompt on four AI video models and save the clips
A Python script that sends one prompt to Wan 3.0, Seedance 2.5, Kling 3 and MiniMax H3 on Sume, polls all four jobs and saves each MP4, with costs.

To compare video models fairly, send the same prompt to each and look at the clips side by side. On Sume that is one Python script: four POST /v1/videos calls that differ only in model and resolution, one polling loop over the four polling_url values, and four downloads from unsigned_urls[0]. The flow follows the video generation docs, read on 2026-10-06.
At 5 seconds and the lowest resolution each model offers for this test, the four jobs cost $0.32 (Wan 3.0, 480p), $1.35 (Seedance 2.5, 480p), $0.70 (Kling 3, 720p, audio off) and $0.32 (MiniMax H3, 480p), about $2.69 in all. Prices are provider list times 1.25, rounded up to the cent per job, from the Video Router docs.
The script
It submits all four jobs first, so they run in parallel on the provider side, then polls every 30 seconds, which is the interval the docs suggest. Set SUME_API_KEY first.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A ceramic mug on a wooden desk, steam rising, soft morning light"
PLAN = {"wan-3.0": "480p", "seedance-2.5": "480p",
"kling-3": "720p", "minimax-h3": "480p"}
pending = {}
for model, res in PLAN.items():
r = requests.post("https://api.sume.com/v1/videos", headers=H, timeout=60,
json={"model": model, "prompt": PROMPT, "duration": 5,
"resolution": res, "aspect_ratio": "16:9"})
if r.status_code != 202:
print(model, "rejected:", r.text)
continue
pending[model] = r.json()["polling_url"]
while pending:
time.sleep(30)
for model, url in list(pending.items()):
s = requests.get(url, headers=H, timeout=30).json()
if s["status"] == "completed":
data = requests.get(s["unsigned_urls"][0], headers=H, timeout=120)
open(f"{model}.mp4", "wb").write(data.content)
del pending[model]
elif s["status"] in ("failed", "cancelled"):
print(model, s["status"], s.get("error"))
del pending[model]What each field does in this script
duration: 5 is inside every model's window (Wan 2-30, Seedance 4-30, Kling 4-15, H3 5-15), so none of the four is rejected for length. aspect_ratio: 16:9 is accepted by all four. The resolutions differ because the models do: Kling does not offer 480p, so it runs at 720p, and H3's native high setting is 768p, not 720p.
A non-202 answer is printed and skipped, so one bad model does not stop the others. A 402 here means the workspace balance is below the reserve for that job. A 400 unsupported_capability means a value is outside that model's list, and the error names what is supported.
How to compare the clips honestly
Because the catalog accepts no seed, a rerun on the same model gives a different take. One clip per model is therefore a sample of one. For a real decision, run each prompt two or three times, or run three different prompts that cover your work: a product on a table, a person in motion, and a shot with on-screen text. Judge the failure modes you care about, such as hands, text and camera drift, before you judge beauty.
Also note what differs by design. MiniMax H3 always returns stereo sound. Kling returns sound unless you turn it off, and the audio line is priced separately. Seedance and Wan can return sound too. If you compare clips with the sound on, you are partly comparing audio models.
Before you scale it up
Add an Idempotency-Key header with a hash of the body to make a network retry safe: a repeat returns the original job instead of creating and charging a second one. Add a check against GET /v1/videos/models so a script never submits a value a model does not list; see check a request against the catalog. And when a new model ships, one catalog call shows whether it is listed, so you can add a row to PLAN the day it appears.
Reading the results
Put the four MP4 files in one folder and open them together. Check the first second for composition, the middle for motion, and the last second for drift. Write down which model failed which check, then repeat with a prompt that has a person and one that has text. A pattern over three prompts is a decision; a single clip is an anecdote. When you pick a winner, note its id and the resolution and duration you tested, because those are the values you will pin in production.
Sources
Related posts
More in Developers
- Python TTS cost calculator: Sume job rounding vs per-character rates
A runnable Python function that prices narration lines on Sume (cent rounding, 1-cent minimum) and at flat per-million rates for MAI-Voice-2.1 and Flash.
- R httr2: submit an AI video job, poll it, and save the MP4 (Wan 3.0)
An R script posts a Wan 3.0 job to Sume with httr2, loops on /v1/jobs/{id}/status using next_poll_after_seconds and writes the finished clip to disk.
- Rails Sidekiq job that polls an AI video API with perform_in
A Sidekiq worker reads Sume's /v1/jobs/{id}/status once, then reschedules itself with perform_in using next_poll_after_seconds until the job is terminal.
- Rails Active Job that polls an AI video API: retry_job and wait
An Active Job on Solid Queue reads Sume's job status once and calls retry_job with wait from next_poll_after_seconds until the video job is terminal.
Written by Sume