Vidu Q4, Wan 3.0 and Seedance 2.5: reference inputs compared
Images, videos and audio you can send as references: Vidu Q4 Preview and Wan 3.0 from their pages, wan-3.0 and seedance-2.5 on Sume from the docs.

Vidu Q4 Preview takes up to 15 reference images and up to 3 MP3 audio references, with no video references on its reference-to-video page. Wan 3.0 takes 10 images, 5 videos and 5 audio clips, 20 items in all. On Sume, wan-3.0 keeps Alibaba's 10/5/5 caps and seedance-2.5 accepts 12 references in total (at most 9 images, 3 videos and 3 audio files). Vidu Q4 itself is not listed on Sume.
How do the reference limits line up?
Counts first, then what each model's page says about prompts.
| Input | Vidu Q4 Preview | Wan 3.0 (Alibaba) | On Sume |
|---|---|---|---|
| Images | 1 to 15 | up to 10 | wan-3.0: 10; seedance-2.5: up to 9, 12 total across types |
| Videos | not listed on the page | up to 5, 15 s total | wan-3.0: 5; seedance-2.5: up to 3, 12 total across types |
| Audio | 0 to 3 MP3, 3 to 12 s each | up to 5, 15 s total | wan-3.0: 5; seedance-2.5: up to 3, 12 total across types |
| Prompt reference | @subject syntax not supported on Q4 Preview | Image 1, Video 1, Audio 1 | prompt text; no label syntax in the docs |
Where does each model fit a reference-heavy job?
Pick by the mix of media you hold.
- Many faces or products and no audio: Vidu's 15 images are the highest count here, but you cannot send them to Sume.
- Reference clips and audio together:
wan-3.0is the row with explicit per-type caps and totals. - A mixed set up to 12 files:
seedance-2.5caps images at 9, videos at 3 and audio at 3, with 12 in total. - Clips up to 30 seconds: both Sume rows go to 30 s; Vidu Q4 Preview stops at 16.
How do you check the live limits?
Read supported_input_references for each id from GET /v1/videos/models. The video generation docs say the Seedance 2.x models, Wan 3.0 and MiniMax H3 accept audio and video references, while Gemini Omni Flash 1.1 accepts video but not audio. The counts are enforced at submit and return a 400 naming the model and the field.
Sources
Related posts
More in Comparisons
- YouTube AI label: three auto-detection signals and Sume outputs
YouTube may label video automatically for its GenAI tools, C2PA metadata or internal detection. Sume docs do not mention C2PA, so here is what to check.
- Wan 3.0 vs Seedance 2.5: the keep rate where Wan is cheaper
Wan 3.0 at 720p costs $0.63 per 5 s and Seedance 2.5 costs $2.89 on Sume. Wan wins per kept clip unless its keep rate falls below about 21.8% of Seedance's.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume