AI video model bake-off: one prompt, three models, your own scores

Leaderboards rank models on someone else's prompts. Run your own bake-off: one prompt, three Sume model ids, fixed settings, and a short scoring sheet.

5 min readSume
All posts

The best AI video model for your ads is the one that scores highest on your own prompts, so run a bake-off: send one prompt to three models with identical settings, score the outputs on a sheet, and pick. Hedra's best-models page, read 2026-10-02, ranks models by Elo, with Omni Flash at 1,240 and Kling 3.0/O3 Pro at 1,110; a ranking like that is a starting list, not a verdict on your product shots.

Request fields come from Sume's Video generation docs. Check that each model accepts your duration and ratio first, because limits differ per model.

How do I run the three-model test?

Hold prompt, duration, ratio, and resolution fixed; change only model. The script sends the same body to three ids and prints each polling_url. Poll each until completed, then download unsigned_urls[0]. If a model rejects a value, that is a result too: note it.

import os
import requests

URL = "https://api.sume.com/v1/videos"
HEAD = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A product on a kitchen counter, slow push-in, soft daylight"

for model in ["seedance-2", "seedance-2.5", "wan-3.0"]:
    body = {"model": model, "prompt": PROMPT, "duration": 5,
            "aspect_ratio": "9:16", "resolution": "720p"}
    r = requests.post(URL, headers=HEAD, json=body, timeout=60)
    print(model, r.status_code, r.json().get("polling_url"))

What should I score?

Score each clip 1 to 5 on the things your ad needs, such as product shape, motion, text legibility, and artifacts. Have two people score blind, with file names hiding the model. Average, then look at the spread, not just the winner.

A scoring sheet with one row per model id from the script.
Sume model idProduct shapeMotionArtifacts
seedance-21-51-51-5
seedance-2.51-51-51-5
wan-3.01-51-51-5

Why not trust the leaderboard alone?

An Elo score is an aggregate over someone else's prompts and says nothing about your product shots. Models also differ in length, ratio, and audio, which a single score hides.

How many runs is enough?

Use at least three prompts per model that match real ads. One lucky clip tells you little. Check cost per clip from pricing_skus in GET /v1/videos/models, or from usage.cost on a short test clip, before multiplying runs.

Keep the outputs and scores in a dated folder. When a new model appears on a board, you can add it to the same sheet and compare it on the same prompts, instead of starting over. Rerun the sheet when you change prompts in a big way, because a model that suited the old wording may not suit the new one.

Sources

Related posts

More in Models

All Models posts

Written by Sume