AI video generator comparison table: price, length, sound

An AI video generator comparison on the same columns: billed rate, clip length, resolutions, sound and inputs for 10 video models, dated.

5 min readSume
All posts

A useful AI video generator comparison puts every model on the same columns: what a second costs, how long one clip can run, which resolutions it makes, whether it generates sound, and what it takes as input. The per-second price alone is not enough, because the same rate can buy a different resolution, length or sound.

The table below does that for the 10 video models callable through Sume's POST /v1/videos, in alphabetical order, with no ranking. Limits come from the Video generation docs, the Sume API reference and Sume's catalog code; rates are computed from its pricing code, read 2026-09-28.

How do the AI video generators compare?

One row per model. Inputs: “first and last frame” is image-to-video, and references guide a fresh clip without being frames.

From Video generation, the Sume API reference, and Sume's catalog and pricing code, read 2026-09-28. Rates are provider list × 1.25, before the 5.5% agent fee. List prices change; GET /v1/videos/models has the live values.
Model idBilled rateClip lengthResolutionsSoundInputs
gemini-omni-flash-1.1$0.0375/s 360p, $0.125/s 720p, $0.1875/s 1080p, $0.375/s 4K3–10 s360p, 720p, 1080p, 4K; 16:9 or 9:16AlwaysText; first and last frame; image and video references; video edit
grok-imagine-video-1.5$0.0125/s4–15 s480p, 720pNoOne image, required
kling-3$0.14/s silent, $0.21/s with sound4–15 s720p (read supported_resolutions); 16:9, 9:16, 1:1OptionalText; first and last frame
minimax-h3$0.0625/s 480p, $0.075/s 768p5–15 s480p, 768p; 2K and 4K upscales priced if requestedAlwaysText; first and last frame; image, video and audio references
minimax-h3-max$0.0625/s 480p, $0.10/s 768p, $0.20/s 1080p5–15 s480p, 768p, 1080pAlwaysText; first and last frame; image, video and audio references
seedance-2$0.0175 per 1,000 tokens (about $0.38/s at 720p 16:9)4–15 s480p, 720p, 1080pOptionalText; first and last frame; image, video and audio references
seedance-2-fast$0.014 per 1,000 tokens (about $0.30/s at 720p 16:9)4–15 s480p, 720p (1080p not in the docs)OptionalText; first and last frame; image, video and audio references
seedance-2-mini$0.00875 per 1,000 tokens (about $0.19/s at 720p 16:9)4–15 s480p, 720p (1080p not in the docs)OptionalText; first and last frame; image, video and audio references
seedance-2.5$0.02675 per 1,000 tokens, $0.02925 at 1080p (about $0.58/s at 720p 16:9)4–30 s480p, 720p, 1080pOptionalText; first and last frame; image, video and audio references
wan-3.0$0.0625/s 480p, $0.125/s 720p, $0.25/s 1080p2–30 s480p, 720p, 1080pOptionalText; first and last frame; image, video and audio references

How do I compare prices when models bill differently?

Convert to the clip you will actually make. Most models bill per second of output at a rate set by resolution; kling-3 sets its rate by sound instead. The Seedance ids bill per 1,000 video tokens, counted from the frame's pixels and the clip's length, so the table also shows a 720p 16:9 second for each.

Then multiply by the length you need, and add the 5.5% agent fee. How much does AI video cost per second? works through 10-second and one-minute examples.

  • Send resolution explicitly. With none, wan-3.0 reserves at its 1080p rate in current code.
  • The two MiniMax models and Gemini Omni Flash 1.1 always make sound, so their rates include it.
  • A failed job's reserve is refunded; see Do failed AI video jobs cost money?

Which AI video generator should I pick?

Pick by the constraint your clip has to meet:

  • One clip over 15 seconds: seedance-2.5 or wan-3.0, both up to 30 seconds.
  • 4K: gemini-omni-flash-1.1 lists 4K among its resolutions; minimax-h3 prices 2K and 4K as upscales of its native 768p.
  • Editing a clip you already have: gemini-omni-flash-1.1, the only model here with a video edit mode.
  • Audio references: the Seedance ids, wan-3.0 and the two MiniMax models.
  • One still animated with no sound: grok-imagine-video-1.5.
  • Veo or Sora: not in Sume's catalog. OpenAI's deprecations page lists the Videos API, sora-2 and sora-2-pro with a shutdown date of 2026-09-24; Sora API alternatives covers the move.

What does this table leave out?

  • sume/auto: Sume picks the model and does not disclose which one ran, so it has no row.
  • Quality, look and render time: no row states them, because none of the sources publishes a comparison. Test your own prompt on the models that pass your constraints.
  • Per-model details such as reference counts and aspect ratios: GET /v1/videos/models lists each model's supported_* fields, and AI video length limits by model covers the length errors.

Sources

Related posts

More in Models

All Models posts

Written by Sume