Vidu Q4 takes 15 reference images; Sume rows cap at 9 or 10

Vidu Q4 accepts 1 to 15 reference images. Sume has no Vidu row: most rows allow 9 images, Wan 3.0 and Omni allow 10. Caps for images, video and audio refs.

4 min readSume
All posts

Vidu Q4 accepts 1 to 15 reference images in reference-to-video, according to its product page (read 2026-10-10). Sume has no Vidu row, and its video rows cap images lower: 9 on most rows, and 10 on Wan 3.0 and Gemini Omni Flash 1.1.

If your Vidu workflow uses 11 to 15 images, you have to trim the set or merge images before you move it to Sume.

Caps by row

The request schema in Sume's API code sets the limits. Rows other than Wan 3.0 and Omni accept at most 9 reference images, 3 reference videos, 3 audio files and 12 references in total. Wan 3.0 accepts 10 images, 5 videos and 5 audio files. Omni accepts 10 images and 3 reference videos of at most 3 seconds each, and no audio.

Wan's constraint text adds that reference videos total at most 15 seconds and reference audio totals at most 15 seconds. Alibaba's Wan 3.0 page (read 2026-10-10) states the same shape: up to 10 images, 5 videos and 5 audio files, and 20 references overall.

Reference caps (Vidu and Alibaba read 2026-10-10; Sume from API schema)
SourceImagesVideosAudio
Vidu Q4 (reference-to-video)1 to 15Not listedUp to 3 voice refs
Alibaba Wan 3.0 pageUp to 10Up to 5Up to 5
Sume wan-3.01055
Sume gemini-omni-flash-1.1103 (each at most 3 s)0
Sume Seedance, MiniMax H3 rows933 (12 total)
Sume kling-3 and Grok row0 / 1 start image00

How to trim a 15-image set

Start by ranking images by how much each one constrains the output. A face, an outfit and a product matter more than five similar backgrounds. Drop duplicates first, then merge. A single board image that holds several small views counts as one image toward the cap, though the model has to read it as a sheet, so test the result before you trust it.

Also remember that the cap is on references, not on total inputs. With 9 images on a Seedance row, you still have 3 slots for video and audio within the 12-reference total.

Cost of reference work on Sume

Sume prices Wan 3.0 per output second, not per reference. A 5-second clip at 720p bills $0.625 with 1 image or with 10. Omni also bills by resolution tier per output second: a 5-second 720p clip bills $0.625 with 10 references. The Seedance price is token-based, so run a test clip and read usage.cost.

In practice the cheap way to iterate is a short Wan clip at 480p: 5 seconds bills $0.3125, so five reference-combination tests bill $1.5625 in total, which is little next to one 5-second Seedance 2.5 clip at 720p ($2.889).

Which row for which job

Pick Wan 3.0 when you want the most references and long output. Pick Omni when you need 4K or a reference clip no longer than 3 seconds. Pick a Seedance row when you want 9 images plus video and audio references in one request. Pick Vidu Q4 itself, outside Sume, when you really do need 15 images in a single call.

The Video Generation docs describe which rows accept which reference types.

What a trimmed set looks like in a request

In POST /v1/videos, references go in input_references as image items, and a model accepts a type only if its supported_input_references includes it. Send the images in the order you want the model to weigh them, since Omni's constraint text refers to references by list position, such as <IMAGE_REF_0> for the first image and <VIDEO_REF_0> for the first video, counted from zero.

If your prompt names the images, reference them the way the row expects. A prompt written for Vidu's numbering may point at the wrong image after you reorder the set.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume