Which AI video models make a 20-second clip in one job on Sume?
Seedance 2.5 (4-30 s) and Wan 3.0 (2-30 s) cover 20 seconds in one job on Sume. Omni Flash 1.1 stops at 10 s, MiniMax H3 at 15 s. Full duration table.

On Sume, two text-to-video models cover a 20-second clip in one generation: seedance-2.5 accepts 4 to 30 seconds and wan-3.0 accepts 2 to 30 seconds. Gemini Omni Flash 1.1 accepts 3 to 10 seconds, MiniMax H3 and H3 Max accept 5 to 15, and each other catalog model has a maximum of 15. Ask GET /v1/videos/models before you submit, because limits differ by model.
The duration table
The table lists what the Sume docs state for supported_durations. Two rows are special: Motion Transfer and Recast take a source video, so their length follows that source and they are not text-to-video.
| Model id | Duration range | Note |
|---|---|---|
seedance-2.5 | 4 to 30 s | 480p, 720p, 1080p |
wan-3.0 | 2 to 30 s | Takes audio and video references |
higgsfield-genjutsu | 4 to 30 s | Motion Transfer: one source video plus 1 to 8 reference images; only when its provider is configured |
h3-max-recast | 5 to 30 s | One source video plus 1 to 4 person photos |
minimax-h3 | 5 to 15 s | Native 480p or 768p |
minimax-h3-max | 5 to 15 s | Faster 768p variant |
gemini-omni-flash-1.1 | 3 to 10 s | 360p to 4K, 16:9 or 9:16 |
| Each other catalog model | Up to 15 s | Read supported_durations |
Why one job beats two clips
A 20-second shot in one generation keeps one camera, one light, and one set of faces. Two 10-second clips joined by a cut can drift: a sleeve changes, a logo moves. If the shot has no cut in it, ask for 20 seconds and keep the continuity. If it is two beats, make two clips on purpose.
If you need more than 30 seconds, you need more than one job. Timeline joins them: its limit is 200 slots and 1800 seconds, at $0.10 per output minute.
Check the request fits
Duration is an integer number of seconds, validated against the model. duration is preferred over the alias duration_seconds, and if you set both they must agree. Ask for a length the model does not list and the request is rejected before it starts, which is cheaper than a failed job.
curl -sS "https://api.sume.com/v1/videos/models" \
-H "Authorization: Bearer $SUME_API_KEY"
# read data[].supported_durations and supported_resolutionsChoosing between them
Pick by what the shot needs. Seedance 2.5 takes 4 to 30 seconds with 480p, 720p, and 1080p. Wan 3.0 starts lower, at 2 seconds, which suits short loops, and also takes audio and video references. If you need a source clip as input, the Motion Transfer and Recast rows above are the choice, and their output length follows the source.
Whichever you pick, read the model's supported_durations from the list endpoint at submit time instead of copying this table into code. The registry can change, and the endpoint is the source of truth.
Pricing is per second
Sume reserves the provider list price times 1.25 at submit, and each model has its own pricing_skus. Seconds drive the bill, so a 30-second job costs three times a 10-second job on the same model and resolution. Check pricing_skus for the model you pick, since rates vary by model and resolution.
Sources
Related posts
More in Media tools
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
- How to extract frames from a video with the Sume API
POST /v1/video-frames returns stills at the times you name from one Sume-hosted clip, as durable images at source size. The call is unbilled.
- How to use Sume's Timeline compose and Timeline audio APIs
Timeline compose puts one still and one video in the same frame as a new MP4. Timeline audio joins or splits Sume-hosted audio into reusable files.
Written by Sume