Longest AI video clip in one request: 30, 15 or 10 seconds by model
Seedance 2.5 and Wan 3.0 reach 30 seconds on Sume; most other rows stop at 15 and Gemini Omni Flash at 10. Ceilings per model, and when to stitch instead.

Two text-to-video models on Sume reach 30 seconds in a single request: Seedance 2.5 (seedance-2.5) and Wan 3.0 (wan-3.0). The Seedance 2.0 family, Kling 3, Grok Imagine Video 1.5 and the MiniMax H3 rows stop at 15 seconds, and Gemini Omni Flash 1.1 stops at 10. Past a model's ceiling you stitch clips on a timeline instead of asking for a longer one.
Ceiling per model
Press coverage of Seedance 2.5 describes it as a 30-second model; the number that binds your request is the one in Sume's catalog, which agrees for the two 30-second rows.
| Model id | Max (s) | Notes |
|---|---|---|
| seedance-2.5 | 30 | 480p, 720p, 1080p |
| wan-3.0 | 30 | 480p, 720p, 1080p |
| kling-3 | 15 | 720p, 1080p |
| seedance-2, seedance-2-fast, seedance-2-mini | 15 | 480p to 1080p |
| minimax-h3, minimax-h3-max | 15 | native 480p and 768p; H3 Max adds 1080p |
| grok-imagine-video-1.5 | 15 | image-to-video only |
| gemini-omni-flash-1.1 | 10 | up to 4K |
What a long clip does and does not give you
A single 30-second request keeps one continuous generation, which suits a walk-through or a single-take product tour. It also costs more per attempt, so a bad take is a more expensive miss. Draft at 480p first, then rerun the winning prompt at 1080p; there is no seed on these rows, so the rerun is a new take, not a replay.
If your idea is a sequence of shots, six 5-second clips usually give you more control than one 30-second prompt, and each can be regenerated alone.
Beyond the ceiling
For more than 30 seconds, generate in segments and join them. Chain by passing the last frame of one clip as the first frame of the next where the model accepts a first frame, and keep aspect ratio and resolution identical across segments so the join does not jump.
Sources
Related posts
More in Models
- Pocket TTS languages: six or seven, and Sume's language field
Kyutai lists six Pocket TTS languages on its model card and blog, seven in the GitHub README. Here is how to read that, and how Sume TTS sets a language.
- Polish, Dutch, Swedish, Turkish text to speech API: Sume pl nl sv tr
Sume's Voices library tags voices pl, nl, sv and tr alongside 12 other languages. What to send for each, and how Eleven v4's list compares.
- Reference audio for AI video: which models accept a voice clip
Seedance 2.5, Wan 3.0 and MiniMax H3 take reference audio on Sume; Kling 3, Gemini Omni Flash and the swap rows do not. Limits and the one-reference rule.
- Omni Flash vs H3 Max for reference-to-video: limits side by side
Gemini Omni Flash 1.1 takes 10 images and 3 short videos; MiniMax H3 Max adds audio refs, 12 files total. Limits, tags and a sample request on Sume.
Written by Sume