Wan 3.0 inputs: text, image, video, audio on Alibaba vs Sume wan-3.0
Alibaba documents text, image, video and audio inputs for Wan 3.0, up to 30 s. Sume's wan-3.0 does text, image plus end frame, and references, 2 to 30 s.

Alibaba's Wan3.0-Video page documents text, image, video and audio inputs, with first-frame and first-and-last-frame image-to-video, reference-based generation, up to 30 seconds, at 480P, 720P and 1080P. Sume's wan-3.0 covers text-to-video, image-to-video with an end frame, and reference-to-video, at 2 to 30 seconds and the same three resolutions.
Input modes side by side
The Alibaba column is from the Model Studio doc, last updated September 28, 2026. The Sume column is from the Video Router and Video Generation docs.
| Mode | Alibaba Wan3.0-Video | Sume wan-3.0 |
|---|---|---|
| Text-to-video | Yes | Yes |
| Image-to-video, first frame | Yes | Yes (frame_images, first_frame) |
| Image-to-video, first and last frame | Yes | Yes (last_frame as well) |
| Reference-based generation | Yes | Yes (input_references) |
| Max length | 30 seconds | 30 seconds (minimum 2) |
| Resolutions | 480P, 720P, 1080P | 480p, 720p, 1080p |
| Audio | Audio listed as an input | Audio generation; audio and video references accepted |
How Sume picks the mode
The Sume Video docs say Wan 3.0 accepts audio and video references, along with the Seedance 2.x and MiniMax H3 models. Gemini Omni Flash 1.1 and h3-max-recast do not take audio references.
Mode is inferred, not declared. If frame_images is present, Sume runs image-to-video. If input_references is present, it runs reference-to-video. With neither, it runs text-to-video. If you send both, frame_images wins.
A first and last frame request
A first-and-last-frame request to wan-3.0 on Sume looks like this. At 720p, 10 seconds bills 10 x $0.125 = $1.25 (a $0.10 fal list rate times 1.25).
{
"model": "wan-3.0",
"prompt": "The door opens and morning light fills the room",
"duration": 10,
"resolution": "720p",
"frame_images": [
{"type": "image_url", "image_url": {"url": "https://example.com/first.png"}, "frame_type": "first_frame"},
{"type": "image_url", "image_url": {"url": "https://example.com/last.png"}, "frame_type": "last_frame"}
]
}Limits of this comparison
The vendor column comes from Alibaba's page, which does not describe any other input type, so this post makes no claim about one. The Sume side is limited to what the Sume docs list. Read supported_input_references and supported_frame_images from GET /v1/videos/models for the live answer, because a model that does not advertise a reference type gets a 400.
Sources
Related posts
More in Models
- What is a full-duplex AI avatar, and when is a rendered clip enough?
A full-duplex avatar listens and talks at once, as in Tavus Griffin-Lite. If your viewers never talk back, a rendered Sume avatar clip does the job.
- Which AI video model for how many seconds: five length bands on Sume
Pick a Sume video model by clip length: 2 s, 3 to 4 s, 5 to 10 s, 11 to 15 s, and 16 to 30 s. Which ids accept each band, from the Video Router docs.
- Which aspect_ratio values do Sume video models list: 9:16 and 16:9
Sume docs list 9:16 and 16:9 for Auto and Gemini Omni Flash 1.1. Other models vary, so read capabilities from the catalog. What to send for TikTok or Reels.
- Which Seedance for a UGC ad: 2.5, 2.0, Fast or Mini on Sume?
Test UGC hooks on Seedance 2.0 Mini or Fast at 480p, then spend on 2.5 only when the ad needs 16-30 s or 1080p. Prices for 15 s vertical clips.
Written by Sume