Which AI video model for a 15-second single take on Sume?
Kling 3.0, Wan 3.0, MiniMax H3 and Seedance 2 reach 15 seconds on Sume; Auto and Grok stop at 10. Seedance 2.5 and Wan 3.0 go to 30. Table of limits.

For a 15-second single take on Sume, pin Kling 3.0, Wan 3.0, MiniMax H3, MiniMax H3 Max or Seedance 2; Auto and Grok Imagine stop at 10 seconds in the panel and Gemini Omni Flash 1.1 stops at 10. Past 15, only Seedance 2.5 and Wan 3.0 go to 30.
Sume's video generation docs put it plainly: limits are not uniform, and every catalog model other than the long-clip rows tops out at 15 seconds. This post turns that into a choice.
What is each model's ceiling?
The numbers below come from the docs and the panel's controls. The panel lists discrete lengths and is a subset of the API, so for the API use supported_durations from the catalog. Durations are whole seconds with no decimals.
| Model | Shortest | Longest | Where listed |
|---|---|---|---|
| Seedance 2.5 | 4 s | 30 s | Video docs |
| Wan 3.0 | 2 s | 30 s | Video docs, panel |
| Seedance 2 | 4 s | 15 s | Video docs |
| Kling 3.0 | 5 s | 15 s | Panel |
| MiniMax H3 | 5 s | 15 s | Video docs, panel |
| MiniMax H3 Max | 5 s | 15 s | Video docs, panel |
| Grok Imagine | 5 s | 10 s | Panel |
| Auto | 3 s | 10 s | Video docs, panel |
| Gemini Omni Flash 1.1 | 3 s | 10 s | Video docs |
Is a 15-second clip one shot or stitched?
A single request returns one clip, so a 15-second take is one generation, not a cut. That is the point of choosing a 15-second-capable model: the camera move and the subject stay continuous. The model still decides how, and a long take is where drift shows up.
Alibaba says Wan 3.0 produces native 30-second clips with a smart duration recommendation and video extension in its Wan3.0 announcement. Sume lists wan-3.0 at 2 to 30 seconds, and the panel offers 2, 5, 6, 8, 10, 12, 15, 20 and 30. Sume asks for an explicit duration; the post on smart duration versus explicit duration covers the difference.
Which should I pick at 15 seconds?
Pick by what else the shot needs. Kling's 3.0 versus 4.0 page says Kling 3.0 supports 3 to 15 seconds and Kling 4.0 up to 30; Sume lists Kling 3.0, so 15 is the ceiling there. If you need 4K, Kling 3.0 is the only panel entry that lists it. If you need stereo audio, MiniMax H3 or H3 Max. If you need references and an end frame together, Wan 3.0 or MiniMax H3 (H3 Max has an end frame but no reference slots in the panel).
- 15 s with 4K in the panel: Kling 3.0.
- 15 s with stereo audio at native 768p: MiniMax H3 or H3 Max.
- 15 s with references and an end frame in the panel: Wan 3.0 or MiniMax H3.
- 15 s through the API with audio and video references: Seedance 2.x.
- 20 to 30 s in one request: Wan 3.0 or Seedance 2.5.
How do I request it?
Send the model and a whole-second duration; the job is asynchronous, so poll or use a callback. Idempotency-Key makes a retry return the original job.
Cost is reserved on submit at provider list times 1.25, so a 15-second clip reserves more than a 5-second one; read usage.cost on the completed job.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: take15-001" \
-d '{"model":"minimax-h3","prompt":"One continuous dolly through a market at dusk","duration":15,"resolution":"768p"}'What should I check before a batch?
A long single take is more likely to drift than a short one, so a 15-second clip is worth one 5-second test of the same prompt and frames first.
Submit one short clip first, open the finished job, and compare what you asked for with what came back. Use an Idempotency-Key on each attempt, because a replay with the same key returns the original job instead of creating and billing a second one. Only then queue the rest.
Sume reserves provider list times 1.25 when a job is submitted, and the poll response's usage.cost is the billable amount. Treat that field, not a panel estimate, as the number to budget with.
What does Sume not do?
Sume does not extend a clip past its model's ceiling inside one request, and Auto does not return more than 10 seconds. For longer sequences, make clips and join them on a timeline, as the post on 30-second AI video in one take versus stitched clips describes.
Sources
Related posts
More in Models
- Does H3 Max Recast keep the original audio? What to check on Sume
fal says H3 Max Recast preserves the source audio. Sume rejects generate_audio and audio references on it. Confirm a result has sound with video inspect.
- flux-2-pro-preview vs flux-2-pro: which id does Sume send?
BFL has flux-2-pro-preview (latest) and flux-2-pro (fixed snapshot). Sume lists black-forest-labs/flux.2-pro and does not let you pick the BFL endpoint.
- gemini-omni-1.1-flash vs gemini-omni-flash-1.1: which id goes where
Google writes the model id gemini-omni-1.1-flash; Sume's catalog id is gemini-omni-flash-1.1. A mapping table, the old preview id, and how to avoid a typo.
- Gemini Omni audio reference: unsupported; Sume models that take audio
Google says Omni's API doesn't accept uploaded audio references. Sume's Omni row has no reference_audio_urls; Seedance 2.x, Wan 3.0 and MiniMax H3 honor audio.
Written by Sume