How long can an AI video be in 2026? Max duration by model
Vendor pages list 20 to 40 seconds per generation in 2026. Sume's catalog ids accept 15 s for most, 30 s for three, and 10 s for Omni Flash 1.1.

In 2026 a single AI video generation runs from about 10 seconds to 30 seconds, and one vendor reaches 40 seconds by extension. Sume's catalog follows its own per-model limits: 30 seconds on seedance-2.5, wan-3.0 and a few specialty ids, 15 seconds on most others, and 10 seconds on Gemini Omni Flash 1.1.
All vendor numbers below come from the vendor's own page, read 2026-10-05. All Sume numbers come from the Sume docs.
Vendor limits
Luma lists clips up to 20 seconds for Ray 3.2. ByteDance Seed lists up to 30 seconds per generation for Seedance 2.5, with multiple rounds of extension. Alibaba's Wan 3.0 page lists native 30-second generation. Google lists Gemini Omni 1.1 Flash scene extension that reads up to 10 seconds of prior context and reaches up to 40 seconds in 10-second steps.
What Sume accepts per id
The Video generation docs say limits differ by model and that you should read supported_durations from GET /v1/videos/models. Sume does not run Ray 3.2. The docs give these ranges:
| Model | Vendor page | Sume id | Sume accepts |
|---|---|---|---|
| Luma Ray 3.2 | up to 20 s | not in catalog | not available |
| Seedance 2.5 | up to 30 s, extendable | seedance-2.5 | 4-30 s |
| Wan 3.0 | native 30 s | wan-3.0 | 2-30 s |
| Gemini Omni 1.1 Flash | up to 40 s by extension | gemini-omni-flash-1.1 | 3-10 s per job |
| Higgsfield Motion Transfer | n/a | higgsfield-genjutsu | 4-30 s |
| H3 Max Recast | n/a | h3-max-recast | 5-30 s |
| MiniMax H3 | n/a | minimax-h3 | 5-15 s |
| MiniMax H3 Max | n/a | minimax-h3-max | 5-15 s |
| Every other catalog id | n/a | see catalog | 15 s maximum |
Reading the table
Two Sume rows need care. higgsfield-genjutsu is Motion Transfer (a source video plus 1-8 reference images) and sits in the catalog only when its provider is configured. h3-max-recast swaps people in a source video, so its duration is the source length. They are editing tools, not text-to-video.
Sume does not offer a scene-extension call like Google's. A longer video on Sume is a chain of jobs joined in Timeline.
How to go longer than one job
Generate clips, extract the last still with video frames, feed it as the next first_frame, and join everything with Timeline 1.0, which renders up to 1800 seconds in one MP4. The 30-second ids need two jobs for a 60-second spot, the 15-second ids need four, and Omni Flash needs six.
Reading the vendor numbers fairly
A vendor's maximum is a product claim, not a guarantee for every plan or resolution. Luma's 20 seconds is listed alongside 1080p. Google's 40 seconds is reached by extension in 10-second steps, so a single generation is shorter. Seedance 2.5's 30 seconds comes with extension rounds listed on top. Where a number is not on a vendor page, this page leaves it out.
For Sume, the numbers are per catalog id and can change, so the check that never goes stale is a call to GET /v1/videos/models.
Pick by the length you need
Up to 10 seconds: any id, including Omni Flash 1.1 at up to 4K. Up to 15 seconds: most ids, including seedance-2 and the MiniMax H3 pair. Up to 30 seconds: seedance-2.5 or wan-3.0 for generation. Beyond 30 seconds in one file: chain clips and join them with Timeline, which renders up to 1800 seconds.
Auto (sume/auto) is limited to 3 to 10 second clips at 16:9 or 9:16 with defaults of 720p and 8 seconds, so do not use Auto when you need a 30-second take.
Cost of going long
Length multiplies cost directly, because Sume bills per output second at the provider list times 1.25. A 30-second job is three times a 10-second job on the same id and resolution. Draft at a low resolution (480p on seedance-2.5, 360p on Omni Flash 1.1) before you commit to a long final render, and use the callback_url option so you are notified instead of polling for several minutes.
The model's maximum is also not the best length. Most social placements use short clips, and a 6 to 8 second shot with a hard cut usually holds attention better than a 30-second take that drifts. Pick the length for the story first, then pick the id that allows it.
Sources
Related posts
More in Comparisons
- Alibaba Live Avatar needs five H800 GPUs: or call a hosted clip API
Alibaba-Quark's open Live Avatar streams audio-driven video at 45 FPS on five H800 GPUs. Compare that hardware with a hosted still-plus-audio clip.
- Anam plans from $12 to $999: cost per live minute vs a rendered one
Anam's plan ladder works out to $0.12 to $0.24 per included live minute. A 60-second Sume avatar clip costs $11.04, so it wins above about 74 viewers.
- Atmee avatar billed by the minute vs Sume reserve and refund
Atmee's LiveKit plugin bills by the minute, capped at 3,600 seconds by default. Sume reserves a clip price upfront and refunds failures. Compared.
- Atmee LiveKit avatar from one portrait vs a Sume photo avatar clip
Atmee turns one portrait into a live LiveKit avatar. Sume turns one photo into a reusable avatar handle and finished clips. Pick by who waits on the face.
Written by Sume