Kling motion control prompt: appearance only, what to write
On Sume's Kling motion control the prompt steers appearance only. The reference video sets all the motion. Write outfit, light and style, not actions.

On Sume's Kling 3.0 motion control, the optional prompt steers appearance only, because the reference video drives all of the motion. Write what the character looks like and how the shot is lit, and leave out what the character does. A prompt that describes an action competes with the reference and gains nothing.
What the field is for
The API documents the prompt as guidance on appearance. The motion comes from motion_video_url, the identity comes from image_url or the avatar, and the prompt fills the space between them: clothing detail, a colour grade, a time of day, a texture.
Write this, not that
The prompt is optional, so the first test is to run without one. Add words only to fix something that you can see in the output.
| Goal | Write | Do not write |
|---|---|---|
| Change the outfit | wearing a navy raincoat with a yellow hood | walks forward in a raincoat |
| Warmer light | soft golden hour light from the left | sun slowly rising |
| Match a style | clean studio look, matte skin | camera pans up |
| Keep the face | same face and hair as the still | smiles and waves |
Why actions backfire
The model receives the motion from the clip. A prompt that says the character raises both arms asks for a different motion than the reference shows, and the result is a blend that satisfies neither. At best the words are ignored. At worst they distort the pose.
Camera movement belongs to the reference too. If you want a different camera, change the reference, or set character_orientation to image so that the framing of the still is kept, as the camera and orientation page explains.
A body to start from
Start with the shortest body that works, add a prompt only if the first result needs a fix, and change one thing at a time. A new prompt is a new body, so use a new Idempotency-Key for it, and run it at a short duration_seconds first, since a 1-second probe costs $0.1575 and a 30-second run costs $4.725.
If a face or an outfit keeps drifting, the cause is usually the still, not the prompt. A sharper, front-facing, well-lit image with the whole body in frame works better than any extra words.
A short test plan
Run three probes at 1 second each, for $0.4725 in total. Probe one has no prompt. Probe two adds a one-line appearance prompt. Probe three adds a second line about the light. Compare the three side by side, and keep the shortest prompt that gives the look that you want. Most jobs need no more than one sentence.
Related posts
More in Developers
- Kling motion control price in usage.billable_amount_usd_micros
The Sume motion control submit response returns usage.billable_amount_usd_micros. Divide by 1,000,000 to log the reserve: 30 s is 4,725,000 micros, or $4.725.
- Kling motion control webhook_url: a Python receiver that verifies
Send mode async plus webhook_url on a Kling motion control submit, then verify x-sume-webhook-signature (HMAC SHA-256 over timestamp.body) in Python.
- Compare 4 image models in one labelled 2x2 sheet with cost per tile
Send one prompt to four Sume image models, paste the results into a 2x2 Pillow sheet, and print each model id with its billed usage.cost on its own tile.
- Last frame of an Omni clip: fixing frame_time_out_of_range
Asking Sume video-frames for t equal to the clip length fails with frame_time_out_of_range. Take the last frame at duration minus a frame, in Python.
Written by Sume