Wan 3.0 reference limits: 10 images, 5 videos and 5 audio tracks
Sume lists Wan 3.0 references as up to 10 images, 5 videos (15 s total) and 5 audio tracks (15 s total). Sume docs do not claim 50 references.

The Sume catalog lists Wan 3.0 references as up to 10 images, up to 5 videos (15 seconds in total, at 16 fps or higher) and up to 5 audio tracks (15 seconds in total) (as of 2026-10-08). The Sume docs do not state a 50-reference limit for any video model, so this post does not claim one.
What Sume lists
The limits below come from the Sume model catalog constraints for Wan 3.0 and from the video docs, which say the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references.
| Model | Image references | Video references | Audio references |
|---|---|---|---|
| Wan 3.0 | up to 10 | up to 5 (15 s total, 16 fps+) | up to 5 (15 s total) |
| Seedance 2.5 | accepted; count limit not in the catalog | accepted | accepted |
| Gemini Omni Flash 1.1 | up to 10 | up to 3, 3 s each | not accepted |
How they are sent
On POST /v1/videos the references go in input_references. On the Video Router they are reference_image_urls, reference_video_urls and reference_audio_urls. A model uses a type only if supported_input_references lists it, so read that field before sending.
If you send both frame_images and input_references, frame_images decides the mode and the request is treated as image-to-video.
Why the audio and video totals matter
The 15 s totals are the part that surprises people. A 30 s Wan clip can use at most 15 s of audio references, so a full 30 s soundtrack does not fit as a reference. Cut the reference to the stretch that carries the voice or rhythm you want matched.
If a number is missing in the catalog, as with Seedance 2.5's counts, do not assume it. Send a small test request and read the error, or read the model entry from GET /v1/video-router/models.
Cost of a reference-heavy run
References do not change the Wan 3.0 price in the grid, which depends on resolution and seconds only. A 30 s clip at 720p is $3.75 whether it has no references or a full set. Check the live catalog if this changes.
A practical set for a 30 s music-led clip: 3 character images, 1 style video of 8 s and 1 audio reference of 12 s. That stays inside every limit above.
Check the number yourself
To confirm the figure behind this post, read the wan-3.0 entry from GET /v1/video-router/models and check that 30 s and 720p are inside its duration_seconds and resolutions. The grid price for that cell is 375 cents ($3.75). It is the provider list price times 1.25, rounded up to the next cent for the job, as of 2026-10-08.
Use an Idempotency-Key on every create call so a retry cannot start a second paid job, then poll the job until it is complete and download the file from the result. Media inputs for Timeline or for references must be hosted files, so import them first.
None of the prices here say anything about picture quality, motion or how well a model follows a prompt. Prices can change when a provider changes its list, which is why every table carries a date. Re-read the catalog before you commit a large batch, and run one test job at the exact setting you plan to use.
Sources
Related posts
More in Models
- What a 13-point Elo gap means: Wan 3.0 vs Seedance 2.5 win chance
Wan 3.0 leads Seedance 2.5 by 13 Elo on the AA text-to-video board. On the standard Elo scale that is a 51.9% win chance per vote. Arithmetic for other gaps.
- When Seedance 2 Mini or Fast is enough, and when to pay for 2.5
A decision guide for Seedance tiers on Sume: what Mini, Fast and 2.5 cost for the same clip, what only 2.5 can do, and which jobs belong on which tier.
- Which image model for thumbnails? Six Sume picks from 3 to 10 cents
Six Sume image models that list 16:9 and 4:5 for thumbnails and covers, from Qwen Image at 3 cents to Nano Banana 2.1 at 10, with the cost of 12 options.
- Which Sume image model after this week's launches: lookup by need
A lookup table: text drafts, photo edits, label text, 4K, 12+ references. Maps Nano Banana 2.1, Imagen 4, GPT Image 2.5 and Hy Image to what Sume lists.
Written by Sume