Score a Sora replacement: five numbers to log for every clip

Before you cut over from Sora to a Sume video model, log five numbers per clip: status, wall time, usage.cost, duration asked and failures. Script included.

5 min readSume
All posts

Compare a Sora replacement with five numbers per clip, not a feeling: final status, wall-clock seconds from submit to terminal, usage.cost, the duration you asked for, and how many attempts failed. All five come from the submit and poll responses Sume already returns (Sume video docs, read 2026-10-06), so the scorecard costs nothing extra to collect.

The OpenAI Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), so you are choosing a new default and the evidence is yours to produce. Run the same twenty prompts through each candidate model and compare the table, then watch the clips.

The five numbers

Scorecard fields and sources, read 2026-10-06
NumberWhere it comes fromWhat it tells you
Statusstatus on the poll responseShare of jobs that end completed
Wall timeYour clock, submit to terminalWhether the model fits your user's patience
Costusage.cost on a completed jobReal billable USD, not the list price
Duration askedThe duration you sentCost per second, with cost divided by this
Failuresfailed jobs and 4xx on submitPrompts a model refuses or cannot do

A runner

It submits one prompt to each model, polls to a terminal state, and prints one row per model. Pick models whose duration range includes the length you test; here 5 seconds fits all three.

import os
import time
import requests

BASE = "https://api.sume.com/v1/videos"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
PROMPT = "A potter shapes a bowl on a wheel, close-up, warm light"

for model in ("gemini-omni-flash-1.1", "wan-3.0", "minimax-h3-max"):
    t0 = time.time()
    r = requests.post(BASE, headers=H, timeout=30, json={
        "model": model, "prompt": PROMPT, "resolution": "720p" if model != "minimax-h3-max" else "768p",
        "duration": 5})
    if not r.ok:
        print(model, "submit", r.status_code)
        continue
    url = r.json()["polling_url"]
    while True:
        time.sleep(15)
        s = requests.get(url, headers=H, timeout=30).json()
        if s["status"] in ("completed", "failed", "cancelled"):
            break
    cost = (s.get("usage") or {}).get("cost")
    print(model, s["status"], round(time.time() - t0), "s", cost)

Reading the result

  • A cheap model that fails one prompt in five costs more per good clip than the number says. Divide total cost by completed clips.
  • Do not mix durations in one comparison. A 10-second clip and a 5-second clip have different costs per second only if you divide.
  • Wall time varies with queue load. Run each model at least three times and keep the median, not the best.
  • A refused prompt is data. Keep the error text; it tells you what to rewrite.

Then watch the clips side by side with the sound on. The numbers narrow the field; your eyes choose. The three-model smoke test is the shorter version of this script, and the 10 percent canary is the next step once a model passes.

Turning five numbers into a decision

A table of five numbers per model is only useful if you decide in advance what a pass looks like. Write the thresholds before you run: for example, at least 90 percent of jobs completed, median wall time under a limit your users can live with, and cost per keeper under a figure your product can carry. Then the result is a yes or a no, not a debate.

Wall time deserves care. It includes queue time, which depends on load and on your plan's concurrency, so a slow median can mean a busy hour and not a slow model. Run at different times of day, and record when each run started so the pattern is visible.

Keep the scorecard script in the repository and run it again when you change the model, the prompt template or the resolution. A comparison you can repeat in ten minutes is a habit; one you did by hand once is an anecdote, and teams that keep the script notice regressions before their users do.

  • Pick the pass thresholds first and write them in the repository next to the script.
  • Run each model on the same prompts, the same durations and the same ratio.
  • Save the raw poll responses for every job, so you can re-score with a different rule later.
  • Add a human score for the clips you would publish, on a fixed scale, because no field in the response measures taste.
  • Date the scorecard; the catalog and the models move.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume