Assign a Sume video model to each old prompt by clip length

A short Python planner that reads a CSV of old prompt lengths, sends 10 seconds or less to Omni, up to 30 seconds to Wan 3.0, and splits longer clips.

5 min readSume
All posts

Clip length decides the model more than taste does. On Sume, gemini-omni-flash-1.1 accepts 3 to 10 seconds, wan-3.0 accepts 2 to 30, and anything longer than 30 seconds has to be split into several jobs. A twelve-line planner can assign all three cases to a whole CSV of old prompts.

OpenAI lists the Sora video models and the Videos API as removed on 2026-09-24 and names no successor, so a prompt library built there needs a new home for every row. Length is the cheapest column to sort by, because it is already in your data.

The windows that matter

The catalog gives each model a supported_durations list, and the repo docs state the ranges in prose. Omni runs 3 to 10 seconds with a default of 8 and always renders native audio. Wan 3.0 runs 2 to 30 seconds at 480p, 720p and 1080p. Seedance 2.5 also reaches 30 seconds but is priced per video token, so this planner leaves it out of the arithmetic.

Both Omni and Wan 720p cost the same per second on paper: the docs list $0.10 for each, and 0.10 x 1.25 = $0.125. So at 720p, length is the only deciding factor, and a short clip goes to Omni purely because its window is tighter and its 16:9 and 9:16 aspects cover most of the work.

Clip length to model and cost at 720p (list x 1.25, read 2026-10-05)
Old clip lengthPlanJobsCost
8 sgemini-omni-flash-1.1, 8 s18 x 0.125 = $1.00
18 swan-3.0, one job118 x 0.125 = $2.25
30 swan-3.0, one job130 x 0.125 = $3.75
65 swan-3.0, 3 jobs of 22, 22, 21 s3$2.75 + $2.75 + $2.63 = $8.13

The planner

The script reads an inline CSV so it runs as written. Swap the string for open(path) when you have a real file. Each job's cost is rounded up to the cent after the 1.25 multiplier, which is how Sume bills, so the 21-second job is 21 x 0.125 = 2.625, which rounds up to $2.63. Splitting spreads the remainder evenly instead of leaving one tiny tail clip.

import csv, io, math
OMNI_720, WAN_720 = 0.125, 0.125   # per second: list x 1.25 (Sume docs)
SRC = "id,seconds\nhero,8\nlaunch,18\nsaga,65\n"

def plan(seconds):
    if 3 <= seconds <= 10:
        return [("gemini-omni-flash-1.1", seconds, OMNI_720)]
    parts = math.ceil(seconds / 30)
    base, extra = divmod(seconds, parts)
    return [("wan-3.0", base + (i < extra), WAN_720) for i in range(parts)]

for row in csv.DictReader(io.StringIO(SRC)):
    jobs = plan(int(row["seconds"]))
    cost = sum(math.ceil(round(s * r, 6) * 100) / 100 for _, s, r in jobs)
    shots = "+".join(f"{m}:{s}s" for m, s, _ in jobs)
    print(f'{row["id"]}: {shots} -> ${cost:.2f}')

Splitting is a creative decision

The planner can cut a 65-second clip into three equal parts, but it cannot make the three parts look continuous. Each part is its own render. If the old prompt described one unbroken shot, rewrite it as three shots with a handoff each can start from. The stored posts on splitting long clips cover that step.

Anything the planner puts on Wan should be tested once against Omni at the same length where both apply, because the two differ in look and sound.

  • Keep the model name in the output, so reviewers see which family made each part.
  • Reject lengths below 2 seconds up front; nothing in the catalog accepts them. Lengths of 2 seconds go to Wan 3.0, because Omni starts at 3.
  • Run the real catalog check from the pre-flight post before you submit.

What to do with the output

Write the plan to a second CSV with one row per job: old id, part number, model, seconds, and the idempotency key. Build the key from the old id, the part number and the model, so a rerun of the planner followed by a submit pass never creates a second paid job for the same part. If you later change the model for a row, the key changes with it, and a fresh job is the correct result.

Then submit in waves, not all at once. The docs describe generation admission and a queue, and a long CSV can exceed your workspace concurrency. The stored post on concurrency waves shows the arithmetic for a hundred-row batch.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume