Which AI video model for how many seconds: five length bands on Sume

Pick a Sume video model by clip length: 2 s, 3 to 4 s, 5 to 10 s, 11 to 15 s, and 16 to 30 s. Which ids accept each band, from the Video Router docs.

5 min readSume
All posts

Pick the clip length first and the model second. Sume's video catalog has different duration ranges, and a request outside a model's range is not accepted. The Video Router docs state the limits in a few sentences, so this page puts them in bands.

The ranges in the docs

  • wan-3.0: 2 to 30 seconds.
  • seedance-2.5: 4 to 30 seconds at 480p, 720p or 1080p.
  • seedance-2: up to 15 seconds, also 1080p.
  • minimax-h3 and minimax-h3-max: 5 to 15 seconds.
  • gemini-omni-flash-1.1: 3 to 10 seconds.
  • h3-max-recast: 5 to 30 seconds, the length of the source video.
  • higgsfield-genjutsu: 4 to 30 seconds, only when its provider is configured.
  • Each other catalog model has a limit of 15 seconds.

The five bands

Read the rows as which ids take that length.

Which ids accept each clip length (Sume docs, read 2026-10-05)
LengthText or image to video idsNote
2 swan-3.0Only model in the docs with a 2 s floor
3 swan-3.0, gemini-omni-flash-1.1Omni starts at 3 s
4 swan-3.0, seedance-2.5, seedance-2 family, gemini-omni-flash-1.1Seedance starts at 4 s
5 to 10 sAll of the above plus minimax-h3, minimax-h3-maxH3 starts at 5 s; Omni ends at 10 s
11 to 15 swan-3.0, seedance-2.5, seedance-2 family, minimax-h3, minimax-h3-maxOmni drops out above 10 s
16 to 30 swan-3.0, seedance-2.5The only two single-job options

Check against the live catalog

The docs tell you to read capabilities from GET /v1/video-router/models, because each model can have a different envelope. GET /v1/videos/models returns supported_durations for the same purpose. If your code picks a model from a fixed table like the one above, a model that changes its range will start returning 400 errors. Reading the list at runtime avoids a redeploy.

If you need more than 30 seconds, one job is not enough. The longest single-job ranges end at 30 seconds, so a 45-second piece is two jobs plus a join.

Length and price together

Length is both a limit and a cost. For the per-second models, the cost is linear in seconds, so a 10-second clip costs twice a 5-second one before rounding. Omni 720p at 5 seconds bills $0.63 and at 10 seconds bills $1.25. Wan 3.0 at 480p bills $0.32 for 5 seconds and $0.63 for 10.

So the order of questions is: how long must the shot be, which ids accept that length, and which of those fits the resolution and the budget. Rank by price only after the length filter, because the cheapest model overall may not accept your length at all.

Edit and recast are different

h3-max-recast is not a length choice. Its duration is the length of the source video, 5 to 30 seconds, because it replaces the people in a source clip with people from 1 to 4 reference photos. gemini-omni-flash-1.1 in edit mode takes no duration for the same reason. Keep those out of a length picker, and route them by task.

Before you ship anything, read the live pages again: the catalog is public, the pricing page is public, and the docs describe the request fields. A blog post is a snapshot. The catalog, the plan grid and the error table are the things that change, so write your code to read them instead of copying numbers from a page, and re-check when a new model is added.

A good habit is a small log line per submit with the model, resolution, duration, estimated cost, job id and the Idempotency-Key you used. When a job misbehaves, those six fields answer most of the questions support will ask, and they let you compare your estimate with usage.cost and the usage ledger without re-running anything.

Sources

Related posts

More in Models

All Models posts

Written by Sume