Kling 4.0 Omni Reference takes 15 references; Sume limits compared
fal says Kling 4.0 Omni Reference accepts up to 15 references. On Sume, kling-3 takes none; Wan 3.0 and Gemini Omni Flash 1.1 publish their own limits.

fal's Kling 4.0 explainer says Omni Reference accepts up to 15 references and the model takes up to 10 keyframes. Sume's kling-3 takes no reference images or videos. For reference-to-video on Sume use wan-3.0 (up to 10 images, 5 videos and 5 audio files) or gemini-omni-flash-1.1 (up to 10 images and 3 videos).
Reference limits (read 2026-10-04)
Kling 4.0 is not in the Sume catalog and its API is not yet published in the sources I read, so the first row is context only.
| Model | Reference images | Reference videos | Reference audio |
|---|---|---|---|
| Kling 4.0 Omni Reference (fal, citing Kling) | Up to 15 references in total | Counted within the 15 | Not stated |
| kling-3 (Sume) | None | None | None |
| wan-3.0 (Sume) | Up to 10 | Up to 5, 15 s total, at least 16 fps | Up to 5, 15 s total |
| gemini-omni-flash-1.1 (Sume) | Up to 10 | Up to 3, each up to 3 s | Not accepted |
Send references on Sume
References go in input_references. Each entry is an image object. If you also send frame_images, Sume treats the request as image-to-video and the references are not used as guidance.
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-3.0",
"prompt": "The two products rotate on a marble table",
"duration": 10,
"resolution": "720p",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/a.png"}},
{"type": "image_url", "image_url": {"url": "https://example.com/b.png"}}
]
}'Check the live limits
Sume's per-model limits are in the catalog, and supported_input_references lists which types a model accepts. Read it before you build a reference workflow.
If you need more than 10 images
Neither Wan 3.0 nor Gemini Omni Flash 1.1 takes more than 10 reference images on Sume. Pick the most informative ten, or split the scene into shots and render them separately, then join them on the Timeline endpoint.
Sources
Related posts
More in Models
- Lip sync with a hand over the mouth: sync-3 vs Fabric on Sume
Sync Labs says sync-3 handles obstructions on faces. Sume's lip-sync routes start from a still plus audio. What each takes as input and what is promised.
- Luma's 2026 timeline: Ray3.14, Ray3.2, Scenes and Variants
Luma shipped Ray3.14 in January, Ray3.2 in June, Scenes in August and Variants on Oct 1, 2026. What each added, and why to pin model ids.
- Lyria 3.5 blocks artist-voice prompts: how to write briefs that pass
Google's Lyria 3.5 docs note that prompts asking for specific artist voices are blocked. Describe the sound instead, then run it through the Sume Music Router.
- MAI-Transcribe-2-Streaming or a batch STT job: which fits?
MAI-Transcribe-2-Streaming returns first partials in just over 100 ms. Sume's speech-to-text is a batch job up to 10 minutes at $0.01 per minute.
Written by Sume