2-second AI video API: only Wan 3.0 takes it on Sume

A 2-second clip on Sume means Wan 3.0: $0.13 at 480p, $0.25 at 720p, $0.50 at 1080p per job. Which modes work at 2 seconds and where the other models start.

5 min readSume
All posts

If you need a 2-second AI video from Sume, use wan-3.0. It is the only model in the video catalog whose duration range starts at 2 seconds (2 to 30). Gemini Omni Flash 1.1 starts at 3, Seedance, Kling and Grok Imagine Video at 4, and MiniMax H3 at 5.

A 2-second Wan 3.0 clip costs $0.13 at 480p, $0.25 at 720p and $0.50 at 1080p, billed per job.

What does a 2-second Wan 3.0 clip cost at each resolution?

Sume's catalog lists Wan 3.0 at the fal list price from 2026-08-24 and bills list times 1.25, rounded up to the cent for each job. Alibaba's own launch post gives the same list rates for its Model Studio: $0.05, $0.10 and $0.20 per second (read 2026-10-04).

Wan 3.0 on Sume: list and billed price by resolution, read 2026-10-04
ResolutionList per secondBilled per second2 s clip5 s clip
480p$0.05$0.0625$0.13$0.32
720p$0.10$0.125$0.25$0.63
1080p$0.20$0.25$0.50$1.25

Which modes work on a 2-second clip?

The catalog lists Wan 3.0 as text-to-video, image-to-video with an optional end frame, and reference-to-video, with audio. The 2-second floor applies to all of them, because it is a property of the model's duration range and not of one mode.

Two Wan options are not exposed in v1: file_url/web_url inputs and enable_thinking. Aspect ratios are 16:9, 4:3, 1:1, 3:4 and 9:16, with no 21:9.

  • Text-to-video: prompt, duration, resolution, aspect ratio.
  • Image-to-video: frame_images with a first_frame, optionally a last_frame.
  • Reference-to-video: input_references with image, video or audio entries.
  • If both frame_images and input_references are sent, frame_images wins and the job is image-to-video.

What does the request look like?

A minimal request pins the model and sets duration to 2. It reads the key from the environment.

import os, requests

r = requests.post(
    "https://api.sume.com/v1/videos",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    json={
        "model": "wan-3.0",
        "prompt": "A coffee cup lid pops open, steam rises",
        "duration": 2,
        "resolution": "720p",
        "aspect_ratio": "9:16",
    },
)
print(r.status_code, r.json())

What if you ask Auto for 2 seconds?

sume/auto defaults to Gemini Omni Flash 1.1, whose envelope is 3 to 10 seconds. A 2-second Auto request is a 400 and is not rerouted to Wan, so pin wan-3.0 explicitly when the length matters.

A sensible habit for short clips is to draft at 480p and move to 720p or 1080p for the final only after you like the motion.

Does Wan keep the 2-second floor everywhere?

Treat the number as a catalog fact that can change. Read supported_durations from GET /v1/videos/models before you hard-code it.

Sources

Related posts

More in Models

All Models posts

Written by Sume