How to compare AI video models fairly: one prompt, three models, 480p
Submit one prompt to Seedance 2.0 Mini, Wan 3.0 and MiniMax H3 at 480p and 5 seconds on Sume, then judge the clips blind. A runnable Python test harness.

A fair model comparison on Sume holds everything fixed except the model: the same prompt, 480p, 5 seconds and the same aspect ratio. Seedance 2.0 Mini, Wan 3.0 and MiniMax H3 all accept that envelope, so one script can submit three jobs and give you three clips to review without knowing which is which.
Fix the variables
Choose values every candidate accepts: 5 seconds is inside all three models' ranges, 480p is a rung on all three, and 16:9 is accepted by all three. There is no seed on any of these models, so run each prompt a few times if the decision matters, and judge more than one take per model.
The harness
It submits one job per model with its own Idempotency-Key and prints the job ids. Poll each with the usual status endpoint afterwards.
import asyncio
import os
import httpx
MODELS = ["seedance-2-mini", "wan-3.0", "minimax-h3"]
PROMPT = "A barista pours oat milk into a latte, slow push-in, warm light"
async def submit(client: httpx.AsyncClient, model: str) -> tuple[str, str]:
body = {"model": model, "prompt": PROMPT, "resolution": "480p",
"duration": 5, "aspect_ratio": "16:9"}
headers = {"Idempotency-Key": f"compare-001-{model}"}
response = await client.post("/v1/videos", json=body, headers=headers)
response.raise_for_status()
return model, response.json()["id"]
async def main() -> None:
key = os.environ["SUME_API_KEY"]
headers = {"Authorization": f"Bearer {key}"}
async with httpx.AsyncClient(base_url="https://api.sume.com", headers=headers) as client:
for model, job_id in await asyncio.gather(*(submit(client, m) for m in MODELS)):
print(model, job_id)
asyncio.run(main())Judge it blind
Rename the downloaded files to A, B and C and have someone who did not run the job rank them on the things you care about: product fidelity, motion, text, audio. Then rerun the winner at the resolution you will ship. At the catalog's dated Wan list rate of $0.05 per second, a 5-second 480p test is $0.25 at list.
| Setting | Value | Why |
|---|---|---|
| Resolution | 480p | lowest rung shared by all three |
| Duration | 5 s | inside every range |
| Aspect ratio | 16:9 | accepted by all three |
| Idempotency-Key | one per model | retries reuse the job |
Sources
Related posts
More in Developers
- Contract-test Sume API responses against openapi.json (pytest)
Validate recorded Sume responses against the OpenAPI schema with jsonschema, including the OpenAPI 3.0 nullable fix. A tested pytest file and fixtures guide.
- DBOS Python durable workflow for a Sume job: resume after a crash
Submit and poll a Sume image job in a DBOS workflow: step retries, order-derived Idempotency-Key and workflow id, tested with DBOS 3.2.0 on SQLite.
- A dry-run flag for Sume API calls: print the request, skip the spend
Add DRY_RUN to code that calls the Sume API: build the body, key and spend cap, print them, and send nothing. Review a batch before it costs money.
- Dub one Short into 8 languages: Python fan-out and the total cost
Detach and transcribe once, then run one TTS job and one render per language. A Python fan-out and the per-Short bill, from Sume's catalog rates.
Written by Sume