Assign a Sume video model to each old prompt by clip length
A short Python planner that reads a CSV of old prompt lengths, sends 10 seconds or less to Omni, up to 30 seconds to Wan 3.0, and splits longer clips.

Clip length decides the model more than taste does. On Sume, gemini-omni-flash-1.1 accepts 3 to 10 seconds, wan-3.0 accepts 2 to 30, and anything longer than 30 seconds has to be split into several jobs. A twelve-line planner can assign all three cases to a whole CSV of old prompts.
OpenAI lists the Sora video models and the Videos API as removed on 2026-09-24 and names no successor, so a prompt library built there needs a new home for every row. Length is the cheapest column to sort by, because it is already in your data.
The windows that matter
The catalog gives each model a supported_durations list, and the repo docs state the ranges in prose. Omni runs 3 to 10 seconds with a default of 8 and always renders native audio. Wan 3.0 runs 2 to 30 seconds at 480p, 720p and 1080p. Seedance 2.5 also reaches 30 seconds but is priced per video token, so this planner leaves it out of the arithmetic.
Both Omni and Wan 720p cost the same per second on paper: the docs list $0.10 for each, and 0.10 x 1.25 = $0.125. So at 720p, length is the only deciding factor, and a short clip goes to Omni purely because its window is tighter and its 16:9 and 9:16 aspects cover most of the work.
| Old clip length | Plan | Jobs | Cost |
|---|---|---|---|
| 8 s | gemini-omni-flash-1.1, 8 s | 1 | 8 x 0.125 = $1.00 |
| 18 s | wan-3.0, one job | 1 | 18 x 0.125 = $2.25 |
| 30 s | wan-3.0, one job | 1 | 30 x 0.125 = $3.75 |
| 65 s | wan-3.0, 3 jobs of 22, 22, 21 s | 3 | $2.75 + $2.75 + $2.63 = $8.13 |
The planner
The script reads an inline CSV so it runs as written. Swap the string for open(path) when you have a real file. Each job's cost is rounded up to the cent after the 1.25 multiplier, which is how Sume bills, so the 21-second job is 21 x 0.125 = 2.625, which rounds up to $2.63. Splitting spreads the remainder evenly instead of leaving one tiny tail clip.
import csv, io, math
OMNI_720, WAN_720 = 0.125, 0.125 # per second: list x 1.25 (Sume docs)
SRC = "id,seconds\nhero,8\nlaunch,18\nsaga,65\n"
def plan(seconds):
if 3 <= seconds <= 10:
return [("gemini-omni-flash-1.1", seconds, OMNI_720)]
parts = math.ceil(seconds / 30)
base, extra = divmod(seconds, parts)
return [("wan-3.0", base + (i < extra), WAN_720) for i in range(parts)]
for row in csv.DictReader(io.StringIO(SRC)):
jobs = plan(int(row["seconds"]))
cost = sum(math.ceil(round(s * r, 6) * 100) / 100 for _, s, r in jobs)
shots = "+".join(f"{m}:{s}s" for m, s, _ in jobs)
print(f'{row["id"]}: {shots} -> ${cost:.2f}')
Splitting is a creative decision
The planner can cut a 65-second clip into three equal parts, but it cannot make the three parts look continuous. Each part is its own render. If the old prompt described one unbroken shot, rewrite it as three shots with a handoff each can start from. The stored posts on splitting long clips cover that step.
Anything the planner puts on Wan should be tested once against Omni at the same length where both apply, because the two differ in look and sound.
- Keep the model name in the output, so reviewers see which family made each part.
- Reject lengths below 2 seconds up front; nothing in the catalog accepts them. Lengths of 2 seconds go to Wan 3.0, because Omni starts at 3.
- Run the real catalog check from the pre-flight post before you submit.
What to do with the output
Write the plan to a second CSV with one row per job: old id, part number, model, seconds, and the idempotency key. Build the key from the old id, the part number and the model, so a rerun of the planner followed by a submit pass never creates a second paid job for the same part. If you later change the model for a row, the key changes with it, and a fresh job is the correct result.
Then submit in waves, not all at once. The docs describe generation admission and a queue, and a long CSV can exceed your workspace concurrency. The stored post on concurrency waves shows the arithmetic for a hundred-row batch.
Sources
Related posts
More in Developers
- Audio detach errors: unsupported_media_type, source_not_found
Each Sume audio detach refusal code and its one-line fix: off-host URL, other workspace, not a video, empty range, no audio track, source too long.
- automation_generation_spend_cap_exceeded: what a 402 cap error ends
The per-run cap rejects only the one generation that would cross it. current, requested and cap come back in micros so you can size the next request.
- Avatar batch: which limit hits first, writes per minute or the queue?
On Pro, 300 writes/min is far above the 24 accepted jobs (4 running, 20 queued). Queue capacity limits an avatar batch first, so submit in waves of that size.
- Avatar create 400: removed name and file fields and their replacements
Old avatar model-run requests that send name or file return 400 with details.fields listing replacements: avatar_handle and input.image_url. Fix both at once.
Written by Sume