Shortest AI video clip via API: 2 seconds on Wan 3.0 only
Sume's video catalog by minimum duration: Wan 3.0 takes 2 seconds, Omni 3, Seedance and Kling 4, MiniMax H3 5. Plus how to trim below a floor.

The shortest AI video clip you can request through Sume is 2 seconds, and only wan-3.0 accepts it: Wan 3.0 takes 2 to 30 seconds. Every other catalog model has a higher floor, 3 seconds on Gemini Omni Flash 1.1, 4 seconds on the Seedance, Kling and Grok Imagine rows, and 5 seconds on the MiniMax H3 family. If you need less than a model's floor, generate its minimum and cut the clip with Sume's video-trim.
The limits below come from Sume's Video Router and Video generation docs and the OpenAPI snapshot, read on 2026-10-03. Durations are whole seconds only; duration is an integer, so 2.5 is not a valid request.
What is the duration range of each video model?
The request schema allows duration from 2 to 30, but each model narrows that. A request outside a model's range is rejected rather than rounded, so read the range before you build a template around it.
Two rows are special. h3-max-recast and higgsfield-genjutsu take their length from the source video you send, so there is no short-clip choice on those routes at all: the output keeps the source length.
| Model id | Minimum | Maximum | Note |
|---|---|---|---|
wan-3.0 | 2 s | 30 s | Shortest floor in the catalog |
gemini-omni-flash-1.1 | 3 s | 10 s | Default 8 s; native audio always on |
seedance-2.5 | 4 s | 30 s | Longest Seedance clip |
seedance-2, seedance-2-fast, seedance-2-mini | 4 s | 15 s | Same envelope across the 2.0 family |
kling-3 | 4 s | 15 s | 720p and 1080p only |
grok-imagine-video-1.5 | 4 s | 15 s | Image-to-video only |
minimax-h3, minimax-h3-max | 5 s | 15 s | 768p is native |
higgsfield-genjutsu | 4 s | 30 s | Equals the source video length |
h3-max-recast | 5 s | 30 s | Equals the source video length |
How do I request a 2-second clip on Wan 3.0?
Pin wan-3.0 on the Video Router and send duration: 2. Wan 3.0 offers 480p, 720p and 1080p, and its audio is optional through generate_audio, so set that explicitly if you care whether the clip has sound.
The submit returns the usual job envelope. Poll status_url, then read result_url once result_ready is true; the artifact URL is under result.artifacts[].
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: wan-two-second-001" \
-d '{
"model": "wan-3.0",
"prompt": "A paper plane lands on a desk, quick whip pan, natural light",
"resolution": "720p",
"duration": 2,
"aspect_ratio": "16:9",
"generate_audio": false,
"mode": "async"
}'How do I get a clip shorter than the model's minimum?
Generate the model's minimum length, then cut it. Sume's video trim takes one clip hosted on media.sume.com, a start, and exactly one of end or duration, and returns a new MP4. The duration field accepts 0.2 to 900 seconds, so a 1.5-second cut from a 4-second Seedance clip is valid. The public rate in the docs is $0.02 per job, with no provider inference.
Trimming does not make generation cheaper. The model still renders its full minimum, and you pay for that. When a 2-second slot is all you need, Wan 3.0 is the direct route; trimming is for when you want a different model's look. The same idea for a longer clip is in Cut a Seedance 2.5 clip shorter: video-trim or regenerate?.
- Pick
wan-3.0when the slot is 2 or 3 seconds and you do not care which model made it. - Pick
gemini-omni-flash-1.1for a 3-second clip with synced audio on, or when you need 4K. - Generate 4 seconds on a Seedance or Kling row and trim when you want that look in a shorter slot.
- Use
precision: "exact"(the default) for a frame-accurate cut, andaudio: "drop"if the cut should be silent.
Does Sume's auto routing help with short clips?
sume/auto generates 3 to 10 seconds, with 720p and 8 seconds as the defaults, so it cannot reach 2 seconds either. Auto also does not tell you which model ran, which matters if a client wants exactly the Wan 3.0 look. See sume/auto duration limit: 3 to 10 seconds for that ceiling. For a 2-second clip, pin wan-3.0.
Sume prices Wan 3.0 per second by resolution, so a 2-second 720p clip costs less than a 30-second one at the same resolution. Read the live rate from GET /v1/videos/models instead of hard-coding it.
What should I check before I rely on a 2-second clip?
Two things. First, check the audio. generate_audio defaults to the model's own capability, so send it explicitly and listen to the result; a 2-second clip leaves very little room for a bad first half-second. Second, check the opening frames of the finished file with video inspect, which returns stills at timestamps you choose, so you can look at second 0 and second 1.5 without downloading the whole clip.
If the clip is going into a longer cut, put it into a Timeline slot and let the timeline set the on-screen length. A slot can be as short as 0.2 seconds, so you can still show only part of a 2-second clip. A model's minimum is the shortest it will generate, not the shortest you can show.
Sources
Related posts
More in Models
- Silent AI video: generate_audio false or drop the audio after
Seedance, Kling and Wan take generate_audio false; Omni and MiniMax H3 always make sound. How to get a silent clip on Sume, and what it costs.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
- Which Sume video model fits your inputs: text, photo, clip, audio
Match the input you hold to a Sume video model: prompt, first frame, end frame, references, audio sample, or a clip to edit. With the 400s each mix causes.
Written by Sume