'Reference inputs must not exceed 12 total files': per-type caps

Most Sume video models take 9 reference images, 3 videos, 3 audios, 12 files total. Wan 3.0 takes 10, 5 and 5; Gemini Omni takes 10 images and 3 short videos.

4 min readSume
All posts

For most pinned models, Video Router allows at most 9 reference images, 3 reference videos and 3 reference audios, and no more than 12 files in total. Wan 3.0 and Gemini Omni Flash 1.1 have their own caps. The error you see depends on which cap you hit first.

The caps, in the order the validation checks them

The request schema first allows up to 10 images, 5 videos and 5 audios for every model, so that the widest model fits. Then a per-model refine narrows it. Wan 3.0 uses the wide numbers. Gemini Omni uses 10 images and 3 videos of at most 3 seconds each. Everything else, including the Seedance and MiniMax H3 families, uses 9, 3 and 3 with a 12-file total.

Reference caps in the Video Router schema on main (read 2026-10-05)
ModelImagesVideosAudiosTotal
wan-3.01055No separate total in the schema
gemini-omni-flash-1.1103 (each up to 3 s)Not acceptedNo separate total
seedance family, minimax-h3, minimax-h3-max93312
kling-3NoneNoneNoneReferences refused
grok-imagine-video-1.51 (as the source frame)NoneNoneOne image

Messages you will see

The wording is specific enough to tell the cause apart: "reference_image_urls must not exceed 9 files.", "reference_video_urls must not exceed 3 files.", "reference_audio_urls must not exceed 3 files." and "Reference inputs must not exceed 12 total files." The Wan form is "model wan-3.0 accepts at most 5 reference_video_urls." and the Gemini form names the model and the three-second limit.

Two limits that count more than files

The catalog adds duration rules the schema does not check by file count: Wan's reference videos total 15 seconds at 16 fps or more, and MiniMax H3 videos run 2 to 15 seconds each, 15 combined. These are recorded as constraints for the model, so validate durations before you upload.

For minimax-h3 the catalog also notes that the first five reference images are free and each additional image costs $0.08 at list, so a nine-image request has a price effect beyond the clip seconds.

Trim a list to the model's cap

The helper picks the cap for the model and cuts the lists to fit. It keeps the order you gave, so put the most important reference first. It prints the result and sends nothing.

CAPS = {"wan-3.0": (10, 5, 5, None), "gemini-omni-flash-1.1": (10, 3, 0, None)}
DEFAULT = (9, 3, 3, 12)

def fit(model, images, videos, audios):
    ci, cv, ca, total = CAPS.get(model, DEFAULT)
    images, videos, audios = images[:ci], videos[:cv], audios[:ca]
    if total:
        images = images[: max(0, total - len(videos) - len(audios))]
    return images, videos, audios

imgs = [f"https://example.com/i{n}.png" for n in range(10)]
vids = ["https://example.com/v1.mp4"] * 3
auds = ["https://example.com/a1.mp3"] * 3
print([len(x) for x in fit("seedance-2.5", imgs, vids, auds)])

When more is not better

A longer reference list dilutes each item. Start with the two or three images that matter and add only if the result drifts. Large lists also cost more to prepare and, on minimax-h3, can add per-image charges after the fifth, so the cap is a ceiling to respect rather than a target to reach.

Sources

Related posts

More in Models

All Models posts

Written by Sume