Cost per second of AI video on Sume: Omni, Wan and H3 Max priced
A per-second and per-clip price table for the video models Sume lists with a fal list price, after the Sora API ended, with the 1.25 multiple already applied.

On Sume, a 5-second 720p clip costs $0.63 on Gemini Omni Flash 1.1 and $0.63 on Wan 3.0, and $0.50 on MiniMax H3 Max at 768p. Those numbers are the provider list price per second times 1.25, rounded up to the cent, from the Video Router catalog (read 2026-10-06). They replace the per-second line you used to read on the OpenAI Sora bill; the Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06).
Seedance models are priced per video token on the same page, so they do not fit a per-second table. For those, ask the catalog for a quote or read pricing_skus from GET /v1/videos/models.
How the price is built
The Video Router doc lists a provider list price for each model with the date it was read at the provider (for example fal 2026-08-28 for Omni). Sume bills list times 1.25, rounded up to cents, on every row including minimax-h3-max. At submit, Sume reserves that billable amount from the workspace USD balance, and usage.cost on the finished job is the Sume billable amount (Sume video docs, read 2026-10-06).
The doc says each row's billable_formula is the per-model truth. If a number below ever disagrees with the catalog response, the catalog wins.
Per second and per clip
List prices are in the Video Router doc; billable is list times 1.25. Clip totals are rounded up to the cent.
| Model and resolution | List per second | Billable per second | 5 s clip | 10 s clip |
|---|---|---|---|---|
| Omni 1.1, 360p | $0.03 | $0.0375 | $0.19 | $0.38 |
| Omni 1.1, 720p | $0.10 | $0.125 | $0.63 | $1.25 |
| Omni 1.1, 1080p | $0.15 | $0.1875 | $0.94 | $1.88 |
| Omni 1.1, 4K | $0.30 | $0.375 | $1.88 | $3.75 |
| Wan 3.0, 480p | $0.05 | $0.0625 | $0.32 | $0.63 |
| Wan 3.0, 720p | $0.10 | $0.125 | $0.63 | $1.25 |
| Wan 3.0, 1080p | $0.20 | $0.25 | $1.25 | $2.50 |
| MiniMax H3 Max, 480p | $0.05 | $0.0625 | $0.32 | $0.63 |
| MiniMax H3 Max, 768p | $0.08 | $0.10 | $0.50 | $1.00 |
| MiniMax H3 Max, 1080p | $0.16 | $0.20 | $1.00 | $2.00 |
Reading the table
- Length limits differ: Omni takes 3-10 seconds, H3 Max 5-15, Wan 2-30. A 10-second price on Omni is its ceiling, while Wan can go to 30 seconds.
- Omni is the only one of the three that lists 16:9 and 9:16 as its whole ratio set. Check
supported_aspect_ratiosfor the others before you budget vertical work. - H3 Max's 1080p is a latent refinement from native 768p, per the catalog notes, so 768p is the native tier and 1080p is an upscale-like step.
- Native audio is part of the price on Omni and H3 Max; neither has a toggle in the docs, so you cannot save money by turning sound off.
Price a job in code
The same arithmetic in Python, so a budget script and the catalog cannot drift. Rates are per-second billable values from the table; replace them with a read of the catalog if you want them live.
import math
BILLABLE_PER_SECOND = {
("gemini-omni-flash-1.1", "720p"): 0.125,
("wan-3.0", "720p"): 0.125,
("minimax-h3-max", "768p"): 0.10,
}
def clip_cost(model: str, resolution: str, seconds: int) -> float:
rate = BILLABLE_PER_SECOND[(model, resolution)]
cents = math.ceil(round(rate * seconds * 100, 6))
return cents / 100
for key in BILLABLE_PER_SECOND:
print(key, clip_cost(*key, 5), clip_cost(*key, 10))The round(..., 6) guards against float noise turning 62.5 cents into 62.50000001 before the ceiling. For ratios of 0.0375 the same guard keeps 3 seconds at 11.25 cents rather than 11.250000000001. If you want the Sora-era comparison of history and tiers, the earlier price history post lays out what OpenAI charged; this page only states what Sume charges now.
Do not multiply by a quantity and expect an exact match for batches. Each job is rounded up on its own, so ten 5-second Omni clips at 720p cost ten times $0.63, which is $6.30, not 10 times $0.625.
Where the per-second view misleads
A per-second rate is a good way to compare models and a poor way to budget a product. Jobs are rounded up per job, durations have minimums and maximums, and resolution changes the rate. The same model can be cheap per second at 480p and not the right fit for a 3-second sting if its minimum is 5 seconds.
Use the rate to shortlist and the per-job figure to plan. For a shortlist, find the rows whose duration range includes your clip length and whose resolution is the lowest you can ship. For a plan, take the per-job price, multiply by the number of jobs you expect including retakes, and compare it with your balance.
- Shortlist by duration range first; a price is irrelevant if the model cannot make the length.
- Use the lowest resolution that survives your delivery channel, and draft below it.
- Plan with per-job figures and count retakes.
- Re-read the catalog when you plan; the list prices come with the date they were read at the provider.
Sources
Related posts
More in Pricing
- Sume TTS vs MAI-Voice-2.1 at 50k to 5M characters a month
At 1M characters a month Sume TTS bills $47.50, MAI-Voice-2.1 lists $22 and Flash $15. The monthly table for 50k to 5M and when the gap is worth paying.
- TTS API price per 1,000 characters: Sume, MAI, Groq, Lemonfox, Unreal
A per-1,000-character table of list prices for Sume, MAI-Voice-2.1 and Flash, Groq Orpheus, Lemonfox and Unreal Speech, with cost for a 750-character short.
- Transcribe 100,000 two-second voice command recordings: the bill
A two-second clip costs about $0.000333 on Sume STT, so 100,000 cost $33.33. Without duration_seconds each job in flight holds a full $0.01 instead.
- Trim first, then caption: the 60-second note on the caption price
Sume quotes the $0.20 video-captions price for videos up to 60 seconds. Cut the episode with video trim at $0.02 first, then caption only the part you keep.
Written by Sume