Reference videos in AI video on Sume: Wan, H3, H3 Max and Omni limits
How many reference images, videos and audio files each Sume video model accepts: Wan 10/5/5, H3 9/3/3, Omni 10/3 with 3-second clips, Kling none.

Four Video Router models take reference media, and their limits are not interchangeable: wan-3.0 accepts up to 10 images, 5 videos and 5 audio files; minimax-h3 and minimax-h3-max accept 9 images, 3 videos and 3 audio files; and gemini-omni-flash-1.1 accepts 10 images and 3 videos but no audio. kling-3 accepts no references at all. The numbers come from the constraint lines in Sume's catalog, which the Video Router docs tell you to read per model.
If you build one request that you replay across models, these differences are where it breaks first.
What are the exact limits?
Video and audio references carry length rules as well as counts. The table copies them from the catalog.
| Model id | Images | Videos | Audio | Length rules |
|---|---|---|---|---|
| wan-3.0 | up to 10 | up to 5 | up to 5 | videos 15 s total, at least 16 fps; audio 15 s total |
| minimax-h3 | up to 9 | up to 3 | up to 3 | each 2-15 s, combined 15 s; images + videos + audios at most 12 |
| minimax-h3-max | up to 9 | up to 3 | up to 3 | each 2-15 s, combined 15 s; images + videos + audios at most 12 |
| gemini-omni-flash-1.1 | up to 10 | up to 3 | none | each video at most 3 s |
| kling-3 | none | none | none | no reference_*_urls |
Which limit do people trip over?
On the H3 family, audio cannot be your only reference, and the cap of 12 applies to images, videos and audios together, so 9 images plus 3 videos already fills it. On Omni each reference video can be at most 3 seconds, which is much shorter than the 15-second videos Wan takes in total. On Wan the 5-video limit is a count while the 15-second figure is a combined length, so five 4-second clips is already over it.
On H3, reference-to-video also has a price detail: the first 5 reference images are free and each additional image carries an $0.08 list add-on, while on H3 Max the catalog says reference-to-video is billed as output seconds at list times 1.25 and fal's reference-token overage is not reserved. Check the line for the specific id before you budget a 9-image job.
How do I check a request before I send it?
The numbers above fit in a small function that counts what you plan to send. It runs offline and prints the problems, so you find a 6th Wan video before the API does.
LIMITS = {
"wan-3.0": (10, 5, 5),
"minimax-h3": (9, 3, 3),
"minimax-h3-max": (9, 3, 3),
"gemini-omni-flash-1.1": (10, 3, 0),
"kling-3": (0, 0, 0),
}
def problems(model, images, videos, audios):
mi, mv, ma = LIMITS[model]
out = []
for n, cap, kind in ((images, mi, "image"), (videos, mv, "video"), (audios, ma, "audio")):
if n > cap:
out.append("%s: %d %s refs, limit %d" % (model, n, kind, cap))
if model.startswith("minimax-h3") and images + videos + audios > 12:
out.append("%s: more than 12 refs in total" % model)
if model.startswith("minimax-h3") and audios and not (images or videos):
out.append("%s: audio cannot be the only reference" % model)
return out
print(problems("wan-3.0", 4, 6, 0))
print(problems("minimax-h3-max", 0, 0, 2))What about the edit and motion-transfer rows?
Two catalog rows work from a source video rather than from references, and they are separate from the table above. gemini-omni-flash-1.1 supports video_to_video: send a prompt and one video_url, and the model edits the clip. The docs say video_url is the edit source, not a reference, and cannot be combined with image_url, end_image_url or reference_*_urls. In edit mode resolution is optional with a default of 720p, aspect_ratio is rejected, and a duration is only a hint for the reserve.
higgsfield-genjutsu (Motion Transfer) takes one source video and 1 to 8 reference images, at 480p or 720p, for 4 to 30 seconds, and the docs say it is listed only when its provider is configured. h3-max-recast swaps the people in a source video for 1 to 4 reference photos, one photo per person. Neither is a general reference-to-video model, so they sit outside the limits table.
How are the limits enforced?
Counts and lengths are limits Sume documents, and some are enforced by the provider. The catalog notes the Omni rule that each reference video must be at most three seconds as enforced by fal and documented by Sume, so a request that breaks it can fail at the provider rather than at admission. Treat the table as the contract and test your longest and heaviest combination once before a batch.
Where should I send what?
If you hold a library of up to 10 product shots and a few short clips, Wan 3.0 has the widest limits and the 30-second ceiling. If your references are a few photos and a voice track, H3 Max takes them natively with stereo sound in the output. If the clip you hold is a 3-second move you want continued, Omni is the model with the 3-second rule. Anything with no references in the plan can go to the cheaper Kling 3 row.
Sources
Related posts
More in Models
- Omni Flash vs H3 Max for reference-to-video: limits side by side
Gemini Omni Flash 1.1 takes 10 images and 3 short videos; MiniMax H3 Max adds audio refs, 12 files total. Limits, tags and a sample request on Sume.
- Runway Veo negativePrompt: 1,000 characters, and Sume's options
Runway's API takes an optional negativePrompt of up to 1,000 characters on three Veo models. Sume's video request lists no such field, so use positive wording.
- Seedance 2.5 'extend twice': what Sume's 30-second request does
ByteDance says Seedance 2.5 makes up to 30 seconds with two extensions. Sume's seedance-2.5 id takes 4-30 seconds per request and has no extend field.
- Seedance 2.5 weak spots: physics and crowded scenes, a test plan
ByteDance says Seedance 2.5 still struggles with complex physics and multi-subject scenes. A 480p test plan with frame checks you can run on Sume.
Written by Sume