90-second explainer from fifteen 6-second scenes, plus the render

Fifteen 6 s scenes at 720p cost $11.25 on Wan 3.0 or Omni and $52.00 on Seedance 2.5; the 90 s timeline render adds $0.20 (2 billed minutes at $0.10).

5 min readSume
All posts

A 90-second explainer built from fifteen 6-second scenes at 720p costs $11.25 in generation on Wan 3.0 or Gemini Omni Flash 1.1 and $52.002 on Seedance 2.5. The timeline render that joins them adds $0.20, because the render is $0.10 per output minute, rounded up, and 90 s rounds to 2 minutes.

Prices on this page are Sume list prices as of 2026-10-08: the provider list price times 1.25, computed from Sume's pricing package and billed per output second, linear in duration inside each model's valid range. The model ids, ranges and input types come from the Video generation docs and the Video Router docs; the per-second numbers can be cross-checked in the public catalog.

Totals

The render price is a fixed Sume price (timeline render $0.10 per output minute, ceil, minimum 1 minute, maximum 30 minutes) from the public catalog read 2026-10-08. Voice-over is a separate line: text-to-speech is $0.0475 per 1,000 characters.

15 scenes × 6 s at 720p plus a 2-minute render, read 2026-10-08
ModelGeneration 90 sRenderTotal
Wan 3.0$11.25$0.20$11.45
Gemini Omni Flash 1.1$11.25$0.20$11.45
Seedance 2.0$34.02$0.20$34.22
Seedance 2.5$52.002$0.20$52.202

Trimming the render

Rounding up is why 90 s costs the same to render as 120 s. If you can cut to 60 s the render falls to $0.10; the 15 scenes at 4 s each would cost $7.60 on Wan 3.0 in total.

On Omni the audio in each scene is native and always on, so a muted voice-over track needs to be mixed over it in the timeline; on Wan 3.0 you can leave generate_audio at the model default or turn it off if the model supports it, which the generate_audio field of GET /v1/videos/models shows.

Scaling this plan up or down

As a yardstick, the plan above is built on Wan 3.0 at 720p, $0.125 per second. Each extra 5 seconds adds $0.625, a further $10 of budget buys 80 more whole seconds, and the largest single job the model accepts (30 s) holds $3.75 at submit. The shortest one (2 s) holds $0.25.

Those three numbers are enough to rescale the plan without a new table. If the plan doubles, double the totals; if a clip is shortened, subtract the seconds multiplied by the rate; and if the tier changes, swap the rate for the one in the tables above.

  • Per second: $0.125
  • Per 5 s: $0.625
  • Per 10 s: $1.25
  • Per 30 s or the model maximum (30 s): $3.75

Check the model's fields first

Before a run, ask GET /v1/videos/models for the model and compare three fields with your plan: supported_durations (every clip length must be listed), supported_resolutions (the tier must be listed, since MiniMax H3 has 768p and no 1080p, and Omni has 360p and 4K) and supported_aspect_ratios. A request outside those lists fails at validation, so the failure costs nothing, but a plan built on a wrong assumption costs a rewrite.

Submitting

Submit through POST /v1/videos with model, prompt, duration in whole seconds and resolution; read GET /v1/videos/models first, because supported_durations and supported_resolutions differ per model. Sume reserves the price at submit and usage.cost on the poll response is the billable amount, so the figures here are what leaves the balance.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume