Pick the cheapest Sume video model for an 8-second 16:9 clip
A Python picker reads the live Sume video catalog, drops models that cannot do your length, aspect or resolution, and ranks the rest by per-second price.

To find the cheapest Sume model for a given clip, read GET /v1/videos/models, drop every model whose supported_durations, supported_aspect_ratios or supported_resolutions exclude your request, and sort what is left by per-second rate times seconds. For an 8-second 16:9 clip at 768p or better, minimax-h3 at 768p comes out lowest in the table below, at $0.60.
OpenAI's deprecations page says the Sora models and Videos API were removed on 2026-09-24. If your old per-clip budget assumed Sora pricing, rebuild it from the numbers you can read today.
Where the rates come from
Sume bills the provider list price times 1.25, rounded up to the cent for each job. The repo docs list per-second provider prices for gemini-omni-flash-1.1 at $0.03, $0.10, $0.15 and $0.30 for 360p, 720p, 1080p and 4K, for wan-3.0 at $0.05, $0.10 and $0.20 for 480p, 720p and 1080p, for minimax-h3 at $0.05 and $0.06 for 480p and 768p, and for minimax-h3-max at $0.05, $0.08 and $0.16 for 480p, 768p and 1080p.
Multiply by 1.25: Omni 720p is 0.10 x 1.25 = $0.125 a second, Wan 720p is the same, minimax-h3 768p is 0.06 x 1.25 = $0.075, and minimax-h3-max 768p is 0.08 x 1.25 = $0.10. The Seedance family is priced per 1,000 video tokens, not per second, so the picker leaves it out rather than guess.
Eight seconds, 16:9, 720p or better
Eight seconds fits every model in the picker: Omni accepts 3 to 10, Wan 2 to 30, and both MiniMax models 5 to 15. The price is the rate times 8.
| Model | Resolution | Per second | 8 seconds |
|---|---|---|---|
| minimax-h3 | 768p | $0.075 | $0.60 |
| minimax-h3-max | 768p | $0.10 | $0.80 |
| gemini-omni-flash-1.1 | 720p | $0.125 | $1.00 |
| wan-3.0 | 720p | $0.125 | $1.00 |
| gemini-omni-flash-1.1 | 1080p | $0.1875 | $1.50 |
| minimax-h3-max | 1080p | $0.20 | $1.60 |
| wan-3.0 | 1080p | $0.25 | $2.00 |
The picker
The rates live in a dictionary because the catalog response carries pricing_skus only as opaque per-SKU strings, and the capability fields are what the code filters on. Update the dictionary when the pricing docs change. The script rounds the product to six places before it takes the ceiling, so floating point never adds a cent.
import json, math, os, urllib.request
RATE = {("gemini-omni-flash-1.1", "360p"): .0375, ("gemini-omni-flash-1.1", "720p"): .125,
("gemini-omni-flash-1.1", "1080p"): .1875, ("wan-3.0", "480p"): .0625,
("wan-3.0", "720p"): .125, ("wan-3.0", "1080p"): .25, ("minimax-h3", "480p"): .0625,
("minimax-h3", "768p"): .075, ("minimax-h3-max", "768p"): .10,
("minimax-h3-max", "1080p"): .20}
def catalog():
req = urllib.request.Request("https://api.sume.com/v1/videos/models",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]})
with urllib.request.urlopen(req, timeout=30) as r:
return {m["id"]: m for m in json.load(r)["data"]}
def cheapest(cat, seconds, aspect, resolutions):
rows = []
for (model, res), rate in RATE.items():
m = cat.get(model)
if (m and res in resolutions and res in m["supported_resolutions"]
and seconds in m["supported_durations"]
and aspect in m["supported_aspect_ratios"]):
rows.append((math.ceil(round(rate * seconds, 6) * 100) / 100, model, res))
return sorted(rows)
if __name__ == "__main__":
for cost, model, res in cheapest(catalog(), 8, "16:9", {"720p", "768p", "1080p"}):
print(f"${cost:.2f} {model} {res}")
Cheap is not the only filter
The cheapest row can still be wrong for the job. Omni always renders native audio, and minimax models add stereo audio, so none of these rows is a silent clip. Omni also only does 16:9 and 9:16, while Seedance 2 accepts 21:9, 1:1 and others, so a 1:1 request needs a model the picker does not list.
Treat the output as a shortlist. Render one golden prompt on the top two rows, look at both, and then pin the winner as a config value.
- Change the aspect argument to "9:16" for vertical work and the same filter still applies.
- A 12-second clip removes Omni from the list because its window ends at 10 seconds.
- Cost is per job after the 1.25 multiplier and cent rounding, so batch totals differ slightly from rate times seconds.
Sources
Related posts
More in Pricing
- Cheapest way to start on Sume video: Pro $40 plus a $10 top-up
Video generation starts on the $40 Pro plan, and wallet top-ups run $10 to $1000. A first month is $50 if no usage is included. What the $10 buys, by model.
- Picture-book narration: 24 pages at 200 characters, MAI vs Sume cost
A 24-page picture book at 200 characters a page is 4,800 characters: 7 cents on MAI-Voice-2.1-Flash, 11 cents on MAI-Voice-2.1 and 23 cents on Sume TTS.
- Cost of 1,000 five-second AI avatar replies: Fabric vs H3 Max lip sync
A bank of 1,000 five-second talking replies costs $312.50 to $1,000 on H3 Max lip sync depending on resolution, and $937.50 on Fabric 720p. Here is the math.
- Cost per finished minute of AI video at 1080p, with retakes included
A finished minute is six 10-second clips plus the retakes you discard. At Sume's Omni 1080p rate that is $11.28 with no retakes and $16.92 at one in two.
Written by Sume