Run one video prompt across five Sume models: a Python bake-off

Submit the same prompt to five Sume video models, choose each model's resolution from the catalog, and compare cost and time. Runnable code under 30 lines.

4 min readSume
All posts

To compare video models fairly, send one prompt to each model at the first resolution and shortest duration its catalog entry lists, then read usage.cost from each finished job. The script below does that against Sume's /v1/videos route: it lists models, picks five, takes the lowest advertised resolution, and prints status and cost per model. It reads limits from the catalog instead of assuming them.

Everything about the route comes from the Video Generation docs: GET /v1/videos/models, POST /v1/videos, polling, Idempotency-Key and usage.cost.

The script

It uses plain requests and no async code. Each submit gets its own Idempotency-Key, so rerunning the script with the same tag will not create duplicate paid jobs. It runs real jobs and spends real balance, so start with a short duration.

import os, time, requests

B = "https://api.sume.com/v1"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A paper boat drifts down a rain gutter, macro lens"
TAG = "bakeoff-1"

models = requests.get(f"{B}/videos/models", headers=H, timeout=30).json()["data"]
jobs = []
for m in models[:5]:
    dur = min(m["supported_durations"])
    res = m["supported_resolutions"][0]
    r = requests.post(f"{B}/videos", timeout=30,
        headers={**H, "Idempotency-Key": f"{TAG}-{m['id']}"},
        json={"model": m["id"], "prompt": PROMPT, "duration": dur, "resolution": res})
    if r.ok:
        jobs.append((m["id"], r.json()["polling_url"]))
    else:
        print(m["id"], r.status_code, r.text[:120])
for mid, url in jobs:
    while True:
        s = requests.get(url, headers=H, timeout=30).json()
        if s["status"] in ("completed", "failed", "cancelled"):
            break
        time.sleep(15)
    print(mid, s["status"], s.get("usage", {}).get("cost"))

What to look at

Sort the printed lines by cost, then watch each file. Cost per clip is not cost per second, because the shortest supported duration differs by model: the docs list 2 seconds for wan-3.0, 3 for gemini-omni-flash-1.1, 4 for seedance-2.5 and 5 for minimax-h3. Divide usage.cost by the duration printed if you want a rate. The first resolution in each list is usually the lowest, but the catalog order is not documented as sorted, so sort it if the comparison matters.

How do you keep the comparison fair?

Change one thing at a time. Hold the prompt, the aspect ratio and, where the catalog allows, the duration constant, and only vary the model. The script uses each model's shortest duration to keep spend low, which makes clips of different lengths; if you want equal lengths, take the largest value that appears in every model's supported_durations, and skip the model if there is none.

Judge on the same questions every time: does motion follow the prompt, does the subject stay consistent, is there audio when you asked for it (generate_audio defaults to the model's own capability, so set it explicitly if you compare sound). Write the verdict next to the job id so you can find the clip again.

Because each submit carries an Idempotency-Key, the tag in the script is the one thing to change when you want a fresh run. Reuse the tag and you get the original jobs back instead of new charges, which is useful if the script crashes halfway.

Limits

The first five catalog entries are whichever the endpoint returns, and models with special inputs (such as h3-max-recast, which needs a source video) will fail the text-only request; the script prints the error and moves on. Polling every 15 seconds is fine for a handful of jobs; for more, pass an HTTPS callback_url. Download needs the same Authorization header.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume