Run one video prompt across five Sume models: a Python bake-off
Submit the same prompt to five Sume video models, choose each model's resolution from the catalog, and compare cost and time. Runnable code under 30 lines.

To compare video models fairly, send one prompt to each model at the first resolution and shortest duration its catalog entry lists, then read usage.cost from each finished job. The script below does that against Sume's /v1/videos route: it lists models, picks five, takes the lowest advertised resolution, and prints status and cost per model. It reads limits from the catalog instead of assuming them.
Everything about the route comes from the Video Generation docs: GET /v1/videos/models, POST /v1/videos, polling, Idempotency-Key and usage.cost.
The script
It uses plain requests and no async code. Each submit gets its own Idempotency-Key, so rerunning the script with the same tag will not create duplicate paid jobs. It runs real jobs and spends real balance, so start with a short duration.
import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A paper boat drifts down a rain gutter, macro lens"
TAG = "bakeoff-1"
models = requests.get(f"{B}/videos/models", headers=H, timeout=30).json()["data"]
jobs = []
for m in models[:5]:
dur = min(m["supported_durations"])
res = m["supported_resolutions"][0]
r = requests.post(f"{B}/videos", timeout=30,
headers={**H, "Idempotency-Key": f"{TAG}-{m['id']}"},
json={"model": m["id"], "prompt": PROMPT, "duration": dur, "resolution": res})
if r.ok:
jobs.append((m["id"], r.json()["polling_url"]))
else:
print(m["id"], r.status_code, r.text[:120])
for mid, url in jobs:
while True:
s = requests.get(url, headers=H, timeout=30).json()
if s["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(15)
print(mid, s["status"], s.get("usage", {}).get("cost"))What to look at
Sort the printed lines by cost, then watch each file. Cost per clip is not cost per second, because the shortest supported duration differs by model: the docs list 2 seconds for wan-3.0, 3 for gemini-omni-flash-1.1, 4 for seedance-2.5 and 5 for minimax-h3. Divide usage.cost by the duration printed if you want a rate. The first resolution in each list is usually the lowest, but the catalog order is not documented as sorted, so sort it if the comparison matters.
How do you keep the comparison fair?
Change one thing at a time. Hold the prompt, the aspect ratio and, where the catalog allows, the duration constant, and only vary the model. The script uses each model's shortest duration to keep spend low, which makes clips of different lengths; if you want equal lengths, take the largest value that appears in every model's supported_durations, and skip the model if there is none.
Judge on the same questions every time: does motion follow the prompt, does the subject stay consistent, is there audio when you asked for it (generate_audio defaults to the model's own capability, so set it explicitly if you compare sound). Write the verdict next to the job id so you can find the clip again.
Because each submit carries an Idempotency-Key, the tag in the script is the one thing to change when you want a fresh run. Reuse the tag and you get the original jobs back instead of new charges, which is useful if the script crashes halfway.
Limits
The first five catalog entries are whichever the endpoint returns, and models with special inputs (such as h3-max-recast, which needs a source video) will fail the text-only request; the script prints the error and moves on. Polling every 15 seconds is fine for a handful of jobs; for more, pass an HTTPS callback_url. Download needs the same Authorization header.
Sources
Related posts
More in Use cases
- Yearbook slideshow video: photos, narration and music in one render
Build a 90-second class slideshow: 18 photos on a Timeline 1.0 spine with fades, TTS narration, and a Music Router bed. Costs worked out, with limits.
- Score a multi-scene video with AI music: tempo, key, lead instrument
How to brief Sume Music for a video with several scenes: one consistent score or contrasting cues, with tempo, key and lead-instrument rules from the docs.
- Shoppable videos on Shop: make the vertical product clips with Sume
Shopify's Winter '26 Edition covers shoppable videos on Shop. How to produce a set of vertical product clips for a Shopify store from product stills with Sume.
- Sketch plus face photo to a thumbnail with GPT Image 2.5
Send a rough layout sketch and a portrait as two references and ask GPT Image 2.5 for a 16:9 thumbnail. The prompt, a 1280x720 Sume request and what to verify.
Written by Sume