Wan 3.0 takes 10 images, 5 videos, 5 audio clips: Sume check
Wan 3.0 accepts ten images, five videos and five audio clips per prompt. Sume honors audio and video references on Wan 3.0, but caps vary; read the catalog.

Wan 3.0 accepts up to ten images, five videos and five audio clips per prompt, plus text, PDFs, web pages and PowerPoint files, according to The Decoder. Sume documents that Wan 3.0 honors audio and video references, but does not publish Wan's per-type caps, so read them from the catalog.
The vendor limits
The Decoder describes the input mix of Alibaba's Wan 3.0, which also generates clips up to 30 seconds.
| Input | Limit |
|---|---|
| Images | Up to 10 |
| Videos | Up to 5 |
| Audio clips | Up to 5 |
| Other | Text, PDFs, web pages, PowerPoint files |
What Sume documents
The video generation docs say only models whose supported_input_references lists a type accept that type. Audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, while Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio.
Caps differ per model. As a contrast, Gemini Omni Flash 1.1 takes up to 10 reference images and up to 3 reference videos of 3 seconds each, with no audio references. wan-3.0 itself accepts 2 to 30 seconds. I found no statement in the docs about Wan's exact reference counts, and no PDF or PowerPoint input, so do not send those.
Check before you submit
Read the catalog and look at the model row rather than relying on a vendor page.
- Call
GET /v1/videos/modelsand readsupported_input_referencesforwan-3.0. - Check
supported_durationsandsupported_resolutionsin the same row. - Send media as public HTTPS URLs through
input_references. - Start with one reference of each type, then add more.
Sources
Related posts
More in Models
- Wan 3.0 two-second clips for Reddit video ads
wan-3.0 on Sume accepts 2 to 30 seconds, so a 2 s Reddit video ad is a single request. Platform minimums vary; check them against the catalog first.
- What is FLUX 3? The family map: image, video, audio, action
BFL's docs describe FLUX 3 as one family covering image, video with synchronized audio, audio and action. Which pieces have open weights and which do not.
- Where can I use MAI-Voice-2.1 today, and what Sume offers instead
Microsoft lists its Playground, Copilot Audio Expressions and Foundry; the launch post adds OpenRouter and Vercel. What Sume's text to speech gives you instead.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
Written by Sume