1-second AI video API: nothing on Sume goes below 2 seconds

No Sume video model takes duration 1. The floor is 2 seconds on Wan 3.0; the rest start at 3, 4 or 5. How to get a 1-second beat anyway, with trim and cost.

5 min readSume
All posts

You cannot ask Sume for a 1-second AI video. The shortest duration any catalog model accepts is 2 seconds, and only wan-3.0 goes that low. Every other model starts at 3, 4 or 5 seconds, so a duration: 1 request is refused rather than rounded up.

If what you actually need is a one-second beat, such as a logo sting or a reaction cut, generate the shortest clip that fits and cut it down afterwards. That costs one generation plus a $0.02 trim.

What is the minimum duration per Sume video model?

Limits differ per model, so Sume's docs tell you to read capabilities from the catalog instead of assuming one envelope. The table below is the current catalog (read 2026-10-04).

Shortest accepted duration per Sume video model, read 2026-10-04
Model idShortestLongestNote
wan-3.02 s30 sThe only 2-second floor
gemini-omni-flash-1.13 s10 sEdit mode follows the source clip
seedance-2.54 s30 s480p, 720p, 1080p
seedance-2, seedance-2-fast, seedance-2-mini4 s15 sExplicit picks only
kling-34 s15 s720p and 1080p
grok-imagine-video-1.54 s15 sImage-to-video only
minimax-h3, minimax-h3-max5 s15 sNative stereo audio
h3-max-recast5 s30 sLength follows the source video

What does a 1-second request return?

Sume validates against the catalog and answers with a 400 unsupported_capability instead of silently changing the value. The error lists the values that would have been accepted, so you can read the floor straight from the response.

With model: "sume/auto" the envelope is the Gemini Omni Flash 1.1 one: 3 to 10 seconds, 8 seconds by default. Asking Auto for 2 seconds fails closed and is never rerouted to Wan.

How do you get a 1-second clip anyway?

Generate 2 seconds on Wan 3.0 at 480p, then cut the range you want with Video trim. Trim takes one clip hosted on media.sume.com, a start, and exactly one of end or duration (0.2 to 900 seconds). The price is $0.02 per job, and the source clip is untouched.

Step one is the generation request below. It reads the key from the environment and polls until the job leaves the queue.

import os, time, requests

BASE = "https://api.sume.com/v1"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

job = requests.post(f"{BASE}/videos", headers=H, json={
    "model": "wan-3.0",
    "prompt": "A paper plane lands on a desk, soft daylight",
    "duration": 2,
    "resolution": "480p",
}).json()
print(job["id"], job["status"])

while True:
    s = requests.get(job["polling_url"], headers=H).json()
    if s["status"] in ("completed", "failed", "cancelled"):
        break
    time.sleep(10)
print(s["status"], s.get("unsigned_urls"), s.get("usage"))

What does the 2-second route cost?

Wan 3.0 480p lists at $0.05 per second (fal list 2026-08-24 in the catalog). Sume bills list times 1.25 and rounds each job up to the cent, so a 2-second 480p clip is $0.13 and the same clip at 720p is $0.25. Add the $0.02 trim and the 1-second beat costs about $0.15 at 480p.

Compare that with the shortest clip on the next-cheapest per-second row: a 3-second Gemini Omni Flash 1.1 clip at 360p bills $0.12 and at 720p $0.38. Wan at 480p is the only row that sells a 2-second clip at all.

When should you not trim?

Dialogue or a lip-synced line will not survive a cut in the middle of a word. In that case pick a duration that holds the whole line, and trim only silence at the head or tail.

Before you build a pipeline around the floor, call the models endpoint once and read supported_durations; the numbers in this table are a snapshot of 2026-10-04 and a model can be added or widened later.

Sources

Related posts

More in Models

All Models posts

Written by Sume