Wan 3.0 takes 10 images, 5 videos, 5 audio clips: Sume check

Wan 3.0 accepts ten images, five videos and five audio clips per prompt. Sume honors audio and video references on Wan 3.0, but caps vary; read the catalog.

3 min readSume
All posts

Wan 3.0 accepts up to ten images, five videos and five audio clips per prompt, plus text, PDFs, web pages and PowerPoint files, according to The Decoder. Sume documents that Wan 3.0 honors audio and video references, but does not publish Wan's per-type caps, so read them from the catalog.

The vendor limits

The Decoder describes the input mix of Alibaba's Wan 3.0, which also generates clips up to 30 seconds.

Wan 3.0 inputs per prompt (read 2026-10-03)
InputLimit
ImagesUp to 10
VideosUp to 5
Audio clipsUp to 5
OtherText, PDFs, web pages, PowerPoint files

What Sume documents

The video generation docs say only models whose supported_input_references lists a type accept that type. Audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, while Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio.

Caps differ per model. As a contrast, Gemini Omni Flash 1.1 takes up to 10 reference images and up to 3 reference videos of 3 seconds each, with no audio references. wan-3.0 itself accepts 2 to 30 seconds. I found no statement in the docs about Wan's exact reference counts, and no PDF or PowerPoint input, so do not send those.

Check before you submit

Read the catalog and look at the model row rather than relying on a vendor page.

  • Call GET /v1/videos/models and read supported_input_references for wan-3.0.
  • Check supported_durations and supported_resolutions in the same row.
  • Send media as public HTTPS URLs through input_references.
  • Start with one reference of each type, then add more.

Sources

Related posts

More in Models

All Models posts

Written by Sume