Kling 4.0 takes 15 reference assets; what Sume models accept
Kling's page lists up to 15 combined references for Kling 4.0. On Sume the cap is per model: Wan 10/5/5, H3 9/3/3 (12 total), Omni 10 images and 3 short videos.

Kling 4.0 is described as taking up to 15 combined reference assets, but Kling 4.0 is not a callable id on Sume today, so the closest honest answer is a per-model table of what Sume does accept. The caps differ by model: wan-3.0 takes up to 10 images, 5 videos and 5 audio files; MiniMax H3 and H3 Max take up to 9 images, 3 videos and 3 audio files with at most 12 in total; Gemini Omni Flash 1.1 takes up to 10 images and 3 videos of at most 3 seconds each; the Seedance ids and everything else accepting references share a 12 item combined cap.
The Kling numbers are from kling.ai's Kling 4.0 vs 3.0 page, read on 2026-10-03: up to 10 images, up to 5 videos of 30 seconds combined, up to 7 subjects, 15 assets in all. Sume's limits come from the Video generation docs and the Video Router catalog.
Side by side
The table puts the vendor page next to what a Sume request will accept. Counts are of input_references entries on POST /v1/videos, or the reference_*_urls arrays on the Video Router route.
Two things in it matter most. First, Kling's 7 'subjects' have no direct equivalent: Sume has no subject object, only image, video and audio reference URLs, so a subject becomes one or more reference images. Second, kling-3 is the Kling model you can call and it accepts no references at all.
| Model | Images | Videos | Audio | Combined cap |
|---|---|---|---|---|
| Kling 4.0 (kling.ai) | 10 | 5, 30 s combined | not stated | 15 assets |
| kling-3 on Sume | 0 | 0 | 0 | references rejected |
| wan-3.0 | 10 | 5, 15 s combined, 16 fps or more | 5, 15 s combined | per-type caps |
| minimax-h3 / minimax-h3-max | 9 | 3, 2-15 s each, 15 s combined | 3, 2-15 s each, 15 s combined | 12 |
| gemini-omni-flash-1.1 | 10 | 3, each 3 s or less | none | per-type caps |
| seedance-2.5 and other ids with references | within the combined cap | within the combined cap | within the combined cap | 12 |
What a rejected request looks like
Going over a cap returns HTTP 400 with unsupported_capability, for example 'wan-3.0 accepts at most 5 video input_references', and the message names the field. Nothing is billed. For H3 and H3 Max, audio cannot be the only reference: the request needs an image or a video alongside it.
If you were planning a 15-asset Kling 4.0 request, the realistic port is to rank your assets. Keep the strongest ten for wan-3.0, or the strongest nine images plus three clips for minimax-h3-max, and put the rest into the prompt text.
How to choose
Pick by what the references are for. Product shots and a character sheet are image references: Wan or Omni (10) beat H3 (9). Motion or camera reference needs video: Wan allows five, H3 three, Omni three but only 3 seconds each. A voice or music reference needs audio: Wan, H3 or the Seedance ids.
Whichever you pick, call GET /v1/videos/models first and read supported_input_references so the request type matches the id. Billing is list times 1.25 for every Video Router model.
Sources
Related posts
More in Models
- Kling 3.0 starts at 3 seconds; Sume's kling-3 starts at 4
Kling's own page lists 3-15 seconds for Kling 3.0. The kling-3 id on Sume accepts 4-15, so a 3 second request returns unsupported_capability.
- Kyutai Pocket TTS languages: what it speaks vs hosted Sume TTS
Pocket TTS is a 100M-parameter open model you run yourself. Its README lists seven languages; Sume's TTS is hosted, per character, with a voice id.
- LLM knowledge cutoffs: Claude Jun 2026, GPT-6.1 Sol Apr 2026
Claude 5.5 models list a June 2026 cutoff and GPT-6.1 Sol April 30, 2026. A video agent on trends needs a data tool; Sume ships trending-videos search.
- Longest AI video clip in one request: 30, 15 or 10 seconds by model
Seedance 2.5 and Wan 3.0 reach 30 seconds on Sume; most other rows stop at 15 and Gemini Omni Flash at 10. Ceilings per model, and when to stitch instead.
Written by Sume