Kling 3.0 subject binding limits and what Sume kling-3 takes
Kling's subject binding takes up to 4 images, a 3-8 s clip and a voice sample. Sume's kling-3 takes none of them, so lock identity with a first frame.

What are Kling 3.0's subject binding limits?
On Kling's own page, subject binding accepts up to 4 reference images (front, side, back and detail views), video clips between 3 and 8 seconds, and voice recordings of 5 to 30 seconds. You upload master images to an Element Library, switch on the toggle named "Bind Subject to Enhance Consistency", and then prompt with cinematography terms.
The page also recommends pairing positive direction with negative constraints, giving "morphing features" as an example, and describes multi-shot output of up to 6 cuts in a 15 second sequence while the bound subject stays stable.
Does Sume's kling-3 have a subject binding equivalent?
No. The kling-3 row in the Sume catalog lists reference images, reference videos and reference audio as unsupported, and its constraint reads "no reference_*_urls". There is also no Element Library, no voice sample input and, in the `/v1/videos` parameter table, no negative_prompt field.
What the row does take is a first frame and a last frame. A first frame is the closest honest substitute for a bound subject: the character's face, outfit and framing are pinned by the pixels you send, and the prompt only has to describe the motion.
| Kling subject binding input | Kling limit | On Sume kling-3 |
|---|---|---|
| Reference images | Up to 4 | Not accepted; use frame_images first frame |
| Reference video clip | 3 to 8 seconds | Not accepted |
| Voice recording | 5 to 30 seconds | Not accepted; audio is the optional generate_audio flag |
| Negative prompt | Recommended by the guide | No field; write exclusions as positive wording |
How do I keep a character steady across several Sume clips?
Use the same master image as the first frame of every clip, and keep each clip under the 15 second ceiling. Consistency then comes from the input image rather than a stored subject.
If your plan needs several reference photos of one person or a voice that persists, switch the model. Seedance 2.x rows, Wan 3.0 and the MiniMax rows accept input_references, with per-model caps enforced by the API (for example Wan 3.0 takes at most 10 image references, and the Seedance rows at most 12 references in total). The trade-offs are laid out in the elements comparison.
- One master image per character, even lighting, a clear face, the same file for every clip.
- Describe motion and camera in the prompt, not appearance; the image already carries appearance.
- Check each clip's first and last frames before stitching; drift shows up at the cut points first.
What about the morphing-features negative prompt?
Sume does not forward a negative prompt for video, and v1 rejects non-empty provider.options, so there is no hidden way to pass one. The practical move is to phrase the constraint positively in prompt, for example "the same face and hairstyle throughout, no change in outfit". See video negative prompts for what to expect.
That is weaker than a model-level setting, and Sume makes no promise that the wording will hold. It is a mitigation, not a lock.
What should I test before relying on this?
Run three clips from the same first frame at 720p with different prompts and compare the faces at frame one, the middle and the last frame. If drift is visible, move that character to a model with reference inputs. Each submit reserves at the provider list price times 1.25 and the poll response reports usage.cost as the billable amount, so a three-clip test at the shortest length keeps the cost small and knowable up front.
Sources
Related posts
More in Models
- Kling 3.0 in the Sume Videos panel: 5-15 s choices vs API 4-15 s
The panel offers Kling 3.0 at 5, 6, 8, 10 or 15 seconds and no end frame. The API accepts any whole second from 4 to 15 and a last frame.
- Kyutai Pocket TTS languages: what it speaks vs hosted Sume TTS
Pocket TTS is a 100M-parameter open model you run yourself. Its README lists seven languages; Sume's TTS is hosted, per character, with a voice id.
- Longest AI video clip in one request: 30, 15 or 10 seconds by model
Seedance 2.5 and Wan 3.0 reach 30 seconds on Sume; most other rows stop at 15 and Gemini Omni Flash at 10. Ceilings per model, and when to stitch instead.
- LTX-2 diffusion decoder or convolutional decoder: which to use
LTX-2 ships a diffusion decoder (better quality, more VRAM) and a lighter convolutional one. What the README says, a draft-then-final habit, and hosted jobs.
Written by Sume