Score a Sora replacement: five numbers to log for every clip
Before you cut over from Sora to a Sume video model, log five numbers per clip: status, wall time, usage.cost, duration asked and failures. Script included.

Compare a Sora replacement with five numbers per clip, not a feeling: final status, wall-clock seconds from submit to terminal, usage.cost, the duration you asked for, and how many attempts failed. All five come from the submit and poll responses Sume already returns (Sume video docs, read 2026-10-06), so the scorecard costs nothing extra to collect.
The OpenAI Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), so you are choosing a new default and the evidence is yours to produce. Run the same twenty prompts through each candidate model and compare the table, then watch the clips.
The five numbers
| Number | Where it comes from | What it tells you |
|---|---|---|
| Status | status on the poll response | Share of jobs that end completed |
| Wall time | Your clock, submit to terminal | Whether the model fits your user's patience |
| Cost | usage.cost on a completed job | Real billable USD, not the list price |
| Duration asked | The duration you sent | Cost per second, with cost divided by this |
| Failures | failed jobs and 4xx on submit | Prompts a model refuses or cannot do |
A runner
It submits one prompt to each model, polls to a terminal state, and prints one row per model. Pick models whose duration range includes the length you test; here 5 seconds fits all three.
import os
import time
import requests
BASE = "https://api.sume.com/v1/videos"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A potter shapes a bowl on a wheel, close-up, warm light"
for model in ("gemini-omni-flash-1.1", "wan-3.0", "minimax-h3-max"):
t0 = time.time()
r = requests.post(BASE, headers=H, timeout=30, json={
"model": model, "prompt": PROMPT, "resolution": "720p" if model != "minimax-h3-max" else "768p",
"duration": 5})
if not r.ok:
print(model, "submit", r.status_code)
continue
url = r.json()["polling_url"]
while True:
time.sleep(15)
s = requests.get(url, headers=H, timeout=30).json()
if s["status"] in ("completed", "failed", "cancelled"):
break
cost = (s.get("usage") or {}).get("cost")
print(model, s["status"], round(time.time() - t0), "s", cost)Reading the result
- A cheap model that fails one prompt in five costs more per good clip than the number says. Divide total cost by completed clips.
- Do not mix durations in one comparison. A 10-second clip and a 5-second clip have different costs per second only if you divide.
- Wall time varies with queue load. Run each model at least three times and keep the median, not the best.
- A refused prompt is data. Keep the
errortext; it tells you what to rewrite.
Then watch the clips side by side with the sound on. The numbers narrow the field; your eyes choose. The three-model smoke test is the shorter version of this script, and the 10 percent canary is the next step once a model passes.
Turning five numbers into a decision
A table of five numbers per model is only useful if you decide in advance what a pass looks like. Write the thresholds before you run: for example, at least 90 percent of jobs completed, median wall time under a limit your users can live with, and cost per keeper under a figure your product can carry. Then the result is a yes or a no, not a debate.
Wall time deserves care. It includes queue time, which depends on load and on your plan's concurrency, so a slow median can mean a busy hour and not a slow model. Run at different times of day, and record when each run started so the pattern is visible.
Keep the scorecard script in the repository and run it again when you change the model, the prompt template or the resolution. A comparison you can repeat in ten minutes is a habit; one you did by hand once is an anecdote, and teams that keep the script notice regressions before their users do.
- Pick the pass thresholds first and write them in the repository next to the script.
- Run each model on the same prompts, the same durations and the same ratio.
- Save the raw poll responses for every job, so you can re-score with a different rule later.
- Add a human score for the clips you would publish, on a fixed scale, because no field in the response measures taste.
- Date the scorecard; the catalog and the models move.
Sources
Related posts
More in Developers
- Prove a Sume video retry is safe: same key, same job, one charge
A runnable test for a ported Sora worker: submit twice with one Idempotency-Key on /v1/videos, assert the same job id came back, and cancel before it bills.
- Sume job statuses: three vocabularies and a Python normalizer
Ported Sora code checks one status word. Sume has pending, queued and IN_QUEUE depending on the route. A table, and a normalizer you can run with no network.
- Timeout for AI video jobs: set a deadline in your worker, not HTTP
A Sora-era HTTP timeout of 10 minutes breaks on Sume: sync waits cap at 30 s while jobs run for minutes. Use a client deadline and a poll. Python example.
- Callback or polling for a ported video worker: pick by job count
Replacing a Sora polling worker on Sume: when callback_url beats polling, when polling is enough, and one function that does both.
Written by Sume