Is Sume's kling-3 Kling VIDEO 3.0 or Omni 3.0? Read the catalog row
Sume's catalog names kling-3 'Kling Video v3 Pro' and lists no reference inputs. Kling describes Omni 3.0 as reference-driven, so match by capability.

Which Kling 3.0 model is kling-3 on Sume?
Sume does not label it as either one. The video catalog names the row "Kling Video v3 Pro", and the model card calls it "Kling Video v3 Pro (fal)" with no reference-to-video. What you can verify is the capability list, and it lines up with the way Kling describes its prompt-led VIDEO 3.0 product rather than its reference-driven Omni 3.0 product.
That is a match by capability, not a vendor statement from Sume. Do not write "Omni" in your own copy for this id.
How does Kling separate the two models?
Kling's comparison page says VIDEO 3.0 is "prompt-led", with text-to-video, image-to-video and start and end frames, while Omni 3.0 is "reference-driven", with multi-image, element and video-element references. Both support up to 15 seconds. VIDEO 3.0 has native audio with dialogue in Chinese, English, Japanese, Korean and Spanish; Omni 3.0 adds element voice control.
The page's guidance is to choose VIDEO 3.0 for prompt-led storytelling and multi-shot work, and Omni 3.0 when references, element voice control or strong subject consistency are central.
| Capability | Kling VIDEO 3.0 (per Kling) | Kling Omni 3.0 (per Kling) | Sume kling-3 row |
|---|---|---|---|
| Text-to-video | Yes | Not stated on the page | Yes |
| Start and end frames | Yes | Not stated on the page | Yes, first and last frame |
| Image, video element references | Element reference only | Yes | None; references rejected |
| Element voice control | Not listed | Yes | No voice input |
| Native audio | Yes | Yes | Optional generate_audio |
| Length | Up to 15 s | Up to 15 s | 4 to 15 s |
What does the Sume row let me send?
Text-to-video, image-to-video with a first and optional last frame, optional audio, 720p or 1080p, and 16:9, 9:16 or 1:1. The catalog capabilities show reference_images, reference_videos and reference_audios all false, and the row's constraint is "no reference_*_urls".
Anything on Kling's Omni feature list that depends on a reference, such as video element references or element voice control, has no input on this row. See the Kling 4.0 Omni comparison for the newest reference limits next to Sume.
Which Sume model do I use when I need references?
Use a row that lists reference types. Seedance 2.x, Wan 3.0 and the MiniMax rows accept input_references, and the API enforces per-model caps. If Kling's look is the point, kling-3 with a first frame is the route. If identity across clips is the point, pick a reference-capable model and test it.
One new item to track: Kling's blog posted a Kling 4.0 launch announcement on September 30, 2026, and its comparison page lists a longer 3 to 30 second range and up to 15 reference assets. None of that is a Sume model id today; the catalog lists kling-3.
How do I confirm the row has not changed?
Call GET /v1/videos/models and read the kling-3 entry. It reports supported_durations, supported_input_references, supported_frame_images and generate_audio live. If a future id for another Kling model appears, it will show up there with its own limits; pin the id you tested rather than assuming the name carries over.
Sources
Related posts
More in Models
- Veda sparse attention for MiniMax H3: 6.8x attention, 3.1x clip
Veda's sparse attention keeps 10 percent of attention work for MiniMax H3. Why 6.8x on attention becomes 3.1x per clip, and what it needs to run.
- 9:16 vertical AI video: which Sume models take an aspect ratio
Seedance, Wan 3.0, Kling 3, MiniMax H3 and Gemini Omni Flash take 9:16; Grok Imagine, Genjutsu and H3 Max Recast take no aspect ratio. A per-model table.
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which Sume image models make 2K or 4K output, by model
FLUX 3 Image added 4K; Sume's catalog has two ways to ask for big images, a resolution tier or custom pixels. Which models take which, and the 3840 edge cap.
Written by Sume