Which AI video model for a 15-second single take on Sume?

Kling 3.0, Wan 3.0, MiniMax H3 and Seedance 2 reach 15 seconds on Sume; Auto and Grok stop at 10. Seedance 2.5 and Wan 3.0 go to 30. Table of limits.

5 min readSume
All posts

For a 15-second single take on Sume, pin Kling 3.0, Wan 3.0, MiniMax H3, MiniMax H3 Max or Seedance 2; Auto and Grok Imagine stop at 10 seconds in the panel and Gemini Omni Flash 1.1 stops at 10. Past 15, only Seedance 2.5 and Wan 3.0 go to 30.

Sume's video generation docs put it plainly: limits are not uniform, and every catalog model other than the long-clip rows tops out at 15 seconds. This post turns that into a choice.

What is each model's ceiling?

The numbers below come from the docs and the panel's controls. The panel lists discrete lengths and is a subset of the API, so for the API use supported_durations from the catalog. Durations are whole seconds with no decimals.

Single-clip length per model on Sume (read 2026-10-02)
ModelShortestLongestWhere listed
Seedance 2.54 s30 sVideo docs
Wan 3.02 s30 sVideo docs, panel
Seedance 24 s15 sVideo docs
Kling 3.05 s15 sPanel
MiniMax H35 s15 sVideo docs, panel
MiniMax H3 Max5 s15 sVideo docs, panel
Grok Imagine5 s10 sPanel
Auto3 s10 sVideo docs, panel
Gemini Omni Flash 1.13 s10 sVideo docs

Is a 15-second clip one shot or stitched?

A single request returns one clip, so a 15-second take is one generation, not a cut. That is the point of choosing a 15-second-capable model: the camera move and the subject stay continuous. The model still decides how, and a long take is where drift shows up.

Alibaba says Wan 3.0 produces native 30-second clips with a smart duration recommendation and video extension in its Wan3.0 announcement. Sume lists wan-3.0 at 2 to 30 seconds, and the panel offers 2, 5, 6, 8, 10, 12, 15, 20 and 30. Sume asks for an explicit duration; the post on smart duration versus explicit duration covers the difference.

Which should I pick at 15 seconds?

Pick by what else the shot needs. Kling's 3.0 versus 4.0 page says Kling 3.0 supports 3 to 15 seconds and Kling 4.0 up to 30; Sume lists Kling 3.0, so 15 is the ceiling there. If you need 4K, Kling 3.0 is the only panel entry that lists it. If you need stereo audio, MiniMax H3 or H3 Max. If you need references and an end frame together, Wan 3.0 or MiniMax H3 (H3 Max has an end frame but no reference slots in the panel).

  • 15 s with 4K in the panel: Kling 3.0.
  • 15 s with stereo audio at native 768p: MiniMax H3 or H3 Max.
  • 15 s with references and an end frame in the panel: Wan 3.0 or MiniMax H3.
  • 15 s through the API with audio and video references: Seedance 2.x.
  • 20 to 30 s in one request: Wan 3.0 or Seedance 2.5.

How do I request it?

Send the model and a whole-second duration; the job is asynchronous, so poll or use a callback. Idempotency-Key makes a retry return the original job.

Cost is reserved on submit at provider list times 1.25, so a 15-second clip reserves more than a 5-second one; read usage.cost on the completed job.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: take15-001" \
  -d '{"model":"minimax-h3","prompt":"One continuous dolly through a market at dusk","duration":15,"resolution":"768p"}'

What should I check before a batch?

A long single take is more likely to drift than a short one, so a 15-second clip is worth one 5-second test of the same prompt and frames first.

Submit one short clip first, open the finished job, and compare what you asked for with what came back. Use an Idempotency-Key on each attempt, because a replay with the same key returns the original job instead of creating and billing a second one. Only then queue the rest.

Sume reserves provider list times 1.25 when a job is submitted, and the poll response's usage.cost is the billable amount. Treat that field, not a panel estimate, as the number to budget with.

What does Sume not do?

Sume does not extend a clip past its model's ceiling inside one request, and Auto does not return more than 10 seconds. For longer sequences, make clips and join them on a timeline, as the post on 30-second AI video in one take versus stitched clips describes.

Sources

Related posts

More in Models

All Models posts

Written by Sume