Shortest AI video clip via API: 2 seconds on Wan 3.0 only

Sume's video catalog by minimum duration: Wan 3.0 takes 2 seconds, Omni 3, Seedance and Kling 4, MiniMax H3 5. Plus how to trim below a floor.

5 min readSume
All posts

The shortest AI video clip you can request through Sume is 2 seconds, and only wan-3.0 accepts it: Wan 3.0 takes 2 to 30 seconds. Every other catalog model has a higher floor, 3 seconds on Gemini Omni Flash 1.1, 4 seconds on the Seedance, Kling and Grok Imagine rows, and 5 seconds on the MiniMax H3 family. If you need less than a model's floor, generate its minimum and cut the clip with Sume's video-trim.

The limits below come from Sume's Video Router and Video generation docs and the OpenAPI snapshot, read on 2026-10-03. Durations are whole seconds only; duration is an integer, so 2.5 is not a valid request.

What is the duration range of each video model?

The request schema allows duration from 2 to 30, but each model narrows that. A request outside a model's range is rejected rather than rounded, so read the range before you build a template around it.

Two rows are special. h3-max-recast and higgsfield-genjutsu take their length from the source video you send, so there is no short-clip choice on those routes at all: the output keeps the source length.

Duration limits by Sume video model, from Sume docs and OpenAPI, read 2026-10-03
Model idMinimumMaximumNote
wan-3.02 s30 sShortest floor in the catalog
gemini-omni-flash-1.13 s10 sDefault 8 s; native audio always on
seedance-2.54 s30 sLongest Seedance clip
seedance-2, seedance-2-fast, seedance-2-mini4 s15 sSame envelope across the 2.0 family
kling-34 s15 s720p and 1080p only
grok-imagine-video-1.54 s15 sImage-to-video only
minimax-h3, minimax-h3-max5 s15 s768p is native
higgsfield-genjutsu4 s30 sEquals the source video length
h3-max-recast5 s30 sEquals the source video length

How do I request a 2-second clip on Wan 3.0?

Pin wan-3.0 on the Video Router and send duration: 2. Wan 3.0 offers 480p, 720p and 1080p, and its audio is optional through generate_audio, so set that explicitly if you care whether the clip has sound.

The submit returns the usual job envelope. Poll status_url, then read result_url once result_ready is true; the artifact URL is under result.artifacts[].

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: wan-two-second-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A paper plane lands on a desk, quick whip pan, natural light",
    "resolution": "720p",
    "duration": 2,
    "aspect_ratio": "16:9",
    "generate_audio": false,
    "mode": "async"
  }'

How do I get a clip shorter than the model's minimum?

Generate the model's minimum length, then cut it. Sume's video trim takes one clip hosted on media.sume.com, a start, and exactly one of end or duration, and returns a new MP4. The duration field accepts 0.2 to 900 seconds, so a 1.5-second cut from a 4-second Seedance clip is valid. The public rate in the docs is $0.02 per job, with no provider inference.

Trimming does not make generation cheaper. The model still renders its full minimum, and you pay for that. When a 2-second slot is all you need, Wan 3.0 is the direct route; trimming is for when you want a different model's look. The same idea for a longer clip is in Cut a Seedance 2.5 clip shorter: video-trim or regenerate?.

  • Pick wan-3.0 when the slot is 2 or 3 seconds and you do not care which model made it.
  • Pick gemini-omni-flash-1.1 for a 3-second clip with synced audio on, or when you need 4K.
  • Generate 4 seconds on a Seedance or Kling row and trim when you want that look in a shorter slot.
  • Use precision: "exact" (the default) for a frame-accurate cut, and audio: "drop" if the cut should be silent.

Does Sume's auto routing help with short clips?

sume/auto generates 3 to 10 seconds, with 720p and 8 seconds as the defaults, so it cannot reach 2 seconds either. Auto also does not tell you which model ran, which matters if a client wants exactly the Wan 3.0 look. See sume/auto duration limit: 3 to 10 seconds for that ceiling. For a 2-second clip, pin wan-3.0.

Sume prices Wan 3.0 per second by resolution, so a 2-second 720p clip costs less than a 30-second one at the same resolution. Read the live rate from GET /v1/videos/models instead of hard-coding it.

What should I check before I rely on a 2-second clip?

Two things. First, check the audio. generate_audio defaults to the model's own capability, so send it explicitly and listen to the result; a 2-second clip leaves very little room for a bad first half-second. Second, check the opening frames of the finished file with video inspect, which returns stills at timestamps you choose, so you can look at second 0 and second 1.5 without downloading the whole clip.

If the clip is going into a longer cut, put it into a Timeline slot and let the timeline set the on-screen length. A slot can be as short as 0.2 seconds, so you can still show only part of a 2-second clip. A model's minimum is the shortest it will generate, not the shortest you can show.

Sources

Related posts

More in Models

All Models posts

Written by Sume