Launch-week video test: 6 prompts, 4 models, 5 s, about $29 on Sume

A fixed test plan for any new AI video model: six prompts, four Sume models, 5 seconds at 720p. Costs per model, a Python submit script, and a stop rule.

4 min readSume
All posts

Short answer

Six prompts on four Sume models at 5 seconds and 720p cost about $29.03 in total: $3.75 on Wan 3.0, $3.75 on Gemini Omni Flash 1.1, $4.20 on Kling 3 with audio off and $17.33 on Seedance 2.5. Run the same set every time a new model appears, so each launch is judged against the same yardstick. All rates are Sume prices, which are provider list times 1.25.

The budget

A 5 second clip at 720p is $0.625 on Wan 3.0 ($0.125 a second) and on Omni ($0.125 a second), $0.70 on Kling 3 at $0.14 a second with audio off, and $2.889 on Seedance 2.5 (21,600 tokens a second at $0.0214 per 1,000 list, times 1.25). Six clips of each are the totals below.

Six 5-second 720p clips per model, Sume rates (Sume pricing code, read 2026-10-04)
Model idPer clipSix clips
wan-3.0$0.625$3.75
gemini-omni-flash-1.1$0.625$3.75
kling-3 (audio off)$0.70$4.20
seedance-2.5$2.889$17.33
Total$29.03

Choosing the six prompts

Pick prompts that stress different things: a human face, hands doing a task, text on a sign, fast camera movement, a product close-up, and a wide landscape. Write them once and keep them in a file, unchanged, so a model that improves is visible as a model that improves and not as a prompt that got luckier. Do not tune a prompt per model, or you will measure your prompt skill.

Submit the same prompt to all four

The script submits one prompt to each model with an Idempotency-Key so a retry cannot double-bill. It sets audio off for Kling to match the budget above.

import asyncio, json, os, urllib.request

MODELS = ["wan-3.0", "gemini-omni-flash-1.1", "kling-3", "seedance-2.5"]
PROMPT = "A barista pours latte art, slow push-in"

def submit(model):
    payload = {"model": model, "prompt": PROMPT, "duration": 5,
               "resolution": "720p"}
    if model == "kling-3":
        payload["generate_audio"] = False
    req = urllib.request.Request(
        "https://api.sume.com/v1/videos",
        data=json.dumps(payload).encode(),
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "Content-Type": "application/json",
                 "Idempotency-Key": "launch-week-" + model},
    )
    with urllib.request.urlopen(req) as r:
        return json.load(r)

async def main():
    jobs = await asyncio.gather(*(asyncio.to_thread(submit, m) for m in MODELS))
    for job in jobs:
        print(job["model"], job["id"], job["polling_url"])

asyncio.run(main())

A stop rule

Write the stop rule before you start: for example, stop if Seedance 2.5 is not clearly better on at least four of six prompts, because it costs 4.6 times as much per clip as Wan. When a new model appears, replace the weakest of the four, rerun the six prompts and compare against the same checklist.

Scoring the results

Make a six-by-four grid and score each clip from one to five on the same three questions: is the subject right, is the motion believable, and would you ship it. Total each model, then divide the model's cost by its score to see cost per quality point. A cheap model with a low score can cost more per quality point than a dearer one, and the grid makes that visible in a minute.

Caveats

  • Default audio varies by model; Omni's audio is native, and the Kling figure assumes audio off.
  • The usage.cost on each finished job is the billable amount; use it to replace these estimates.
  • Check the catalog first: an id that is not listed returns 404 model_not_found.

Related posts

More in Use cases

All Use cases posts

Written by Sume