Kling 4.0's 15 references vs Sume: which video model takes the most

Kling 4.0 lists up to 15 references: 10 images and 5 videos. On Sume, Wan 3.0 takes 10 images, 5 videos and 5 audio clips; Omni 10 images and 3 short videos.

5 min readSume
All posts

On Sume today, Wan 3.0 takes the most references: up to 10 images, 5 videos and 5 audio clips, with the videos limited to 15 seconds in total. That is the closest match to Kling 4.0's up to 15 references, of which up to 10 are images and up to 5 are videos with 30 seconds combined, per the fal explainer (read 2026-10-05).

Gemini Omni Flash 1.1 takes up to 10 image references and up to 3 video references, each at most 3 seconds. MiniMax H3 and H3 Max take up to 9 images, 3 videos and 3 audio clips, with at most 12 in total. kling-3 takes no reference media at all.

Where the limits come from

The numbers come from the Sume model notes and the Video Router docs. For Wan 3.0 the notes give at most 10 images, at most 5 videos with at most 15 seconds in total and a minimum of 16 frames per second, and at most 5 audio clips with at most 15 seconds. For Omni the Video Router table gives reference_image_urls up to 10 and reference_video_urls up to 3, each up to 3 seconds.

The video docs say which model accepts which type: the Seedance 2.x models, Wan 3.0, MiniMax H3 and H3 Max accept audio and video references; Omni, Higgsfield Genjutsu and H3 Max Recast accept video references but not audio. Seedance 2.5 lists image, video and audio reference support, but the pages read for this post do not state per-type counts.

Where a number is not stated, this post says so, and it does not guess. Seedance 2.5 is the clearest case: the pricing pages show that a reference video changes the cost, and the model lists reference support, but no count is given in the pages read, so test it with the catalog before you design a plan around it.

Reference limits, side by side

Reference inputs per request, read 2026-10-05
ModelImagesVideosAudio
Kling 4.0 (fal page)up to 10 (15 refs in all)up to 5, 30 s combinednot stated
wan-3.0up to 10up to 5, 15 s totalup to 5, 15 s total
gemini-omni-flash-1.1up to 10up to 3, 3 s eachnone
minimax-h3 / h3-maxup to 9 (12 in all)up to 3up to 3
kling-3nonenonenone

What the limits mean

The limits are not interchangeable. Wan 3.0 has no 21:9 and caps video references at 15 seconds in total, half of what the Kling 4.0 page lists. Omni's 3 second videos are good for a short motion cue, not a full clip. The MiniMax pair is the only one on the list that mixes a 12-item cap with 21:9 and native audio.

For a plan with many references, count them first, then pick the model. A character sheet of six images and one motion clip fits Wan 3.0 and the MiniMax pair, and fits Omni if the clip is 3 seconds or less.

Counting matters, so check a plan before you submit it. The video catalog lists the numbers for each model in supported_input_references, and a check in your code saves a round trip and keeps a rejected request from looking like a model fault.

Which models fit a set

This checks a reference set against the three models and prints which fit. It does not call the API.

Notice that audio references are the one count where Omni drops out. If your plan has a voice clip or a music reference, the choice is between Wan 3.0 and the MiniMax pair, and then 21:9 decides: MiniMax takes it and Wan 3.0 does not. If the plan is images only, any of the three works, and price decides.

LIMITS = {
    "wan-3.0": {"images": 10, "videos": 5, "audios": 5, "total": None},
    "gemini-omni-flash-1.1": {"images": 10, "videos": 3, "audios": 0, "total": None},
    "minimax-h3": {"images": 9, "videos": 3, "audios": 3, "total": 12},
}

def fits(model, images, videos, audios):
    lim = LIMITS[model]
    if images > lim["images"] or videos > lim["videos"] or audios > lim["audios"]:
        return False
    return lim["total"] is None or images + videos + audios <= lim["total"]

want = (8, 3, 2)
for m in LIMITS:
    print(m, fits(m, *want))

Sending them

Send references on POST /v1/videos as input_references, or on the Video Router as reference_image_urls, reference_video_urls and reference_audio_urls, and read supported_input_references from GET /v1/videos/models first. If you send frame_images as well, the docs say that frame_images wins. The jobs docs cover the polling.

On the OpenRouter-compatible route each image reference is an object with type set to image_url and an image_url object holding the url. Keep every reference at a URL that Sume can fetch, and keep the number within the limits above. A request that mixes first and last frames with references will use the frames, so decide which you mean.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume