Kling 3.0 Element Binding: a face close-up, and one still on Sume
Kling says 3.0 Motion Control wants a clear facial close-up for Element Binding. Sume's route takes one still or a ready avatar. What that means for faces.

Kling's guide says that for Video 3.0 Motion Control you should upload clear facial close-ups so Element Binding has enough facial data (read 2026-10-07). Sume's Kling 3.0 Motion Control route has no Element Binding field: it takes exactly one visual source, either a public HTTPS image_url or a ready avatar. So on Sume the face quality is decided by the single still you send.
That is a real limit, and it is easier to plan around than to discover after a failed take. This post says what the Kling guide asks for, what Sume accepts, and how to pick the still.
What Kling's guide says about faces
Kling describes Video 3.0 Motion Control as a feature that uses Element Binding for facial data. The guide tells you to supply clear facial close-ups for Element Binding, and to keep the character's head and body visible in the main image.
| Topic | Kling guide | Sume route |
|---|---|---|
| Face data | Upload clear facial close-ups for Element Binding | No separate close-up field |
| Main image | Head and body clearly visible, not obstructed | One image_url or one avatar |
| Orientation | Matches video (default) or matches image | character_orientation: video or image |
What Sume accepts
The request needs exactly one of image_url, avatar_id, or avatar_handle, plus motion_video_url and duration_seconds from 1 to 30. The body is strict: fields that belong to other surfaces, such as reference image lists, are rejected. The Sume docs do not list an element or face-reference field, so do not expect one.
That leaves two practical choices. Send a still where the face fills a good part of the frame, or use a ready avatar, whose identity still Sume resolves on the server side.
How to pick the single still
A close-up still and a full-body clip can fight each other: Kling's guide asks you to match the body proportion of the image to the motion reference. When in doubt, choose the still that matches the clip.
- Frame the character the way the motion clip frames its performer: full body for a dance, half body for a talking gesture.
- Keep the face large and sharp inside that frame. If the face is tiny, no setting on Sume can add data back.
- Avoid hands or hair covering the face, as Kling's guide asks that the head not be obstructed.
- If the character will return in many clips, save a ready avatar once and reuse its handle instead of re-uploading stills.
What it costs to find out
Sume bills ceil(duration_seconds) times $0.126 times 1.25, which is $0.1575 per second. A 6-second test of a new still reserves $0.945 and is refunded if the job fails. Run one short test per new character before you commit a 30-second take at $4.725.
What Sume does not do
The Sume route runs the Kling 3.0 Standard Motion Control queue and documents no professional-mode switch and no per-subject binding for several characters in one scene. It animates one still.
Sources
Related posts
More in Media tools
- Kling motion control: 3 s minimum on Kling, 1 s field on Sume
Kling says action videos run 3 to 30 seconds. Sume's duration_seconds accepts 1 to 30. Why to keep the reference at 3 seconds or more, and what it costs.
- Kling motion control still: 340 to 3,850 px edges, checked first
Kling's motion control guide sets a 340 px short edge and a 3,850 px long edge for the still. Check both in Python before a Sume job reserves money.
- Meta wants a fixed frame rate: set Timeline output.fps when clips mix
Meta's Facebook feed page requires a fixed frame rate. Sume Timeline takes output.fps 24, 25, 30 or 60, and mixed-rate sources are resampled with judder.
- Bumper and non-skippable ads bill per impression: cut 6, 15, 30 s
Google bills bumper and non-skippable in-stream per impression (Target CPM). Length is a creative call: cut 6, 15, and 30 second versions with video-trim.
Written by Sume