Kling motion control prompt: appearance only, what to write

On Sume's Kling motion control the prompt steers appearance only. The reference video sets all the motion. Write outfit, light and style, not actions.

5 min readSume
All posts

On Sume's Kling 3.0 motion control, the optional prompt steers appearance only, because the reference video drives all of the motion. Write what the character looks like and how the shot is lit, and leave out what the character does. A prompt that describes an action competes with the reference and gains nothing.

What the field is for

The API documents the prompt as guidance on appearance. The motion comes from motion_video_url, the identity comes from image_url or the avatar, and the prompt fills the space between them: clothing detail, a colour grade, a time of day, a texture.

Write this, not that

The prompt is optional, so the first test is to run without one. Add words only to fix something that you can see in the output.

Prompt text for motion control (read 2026-10-05)
GoalWriteDo not write
Change the outfitwearing a navy raincoat with a yellow hoodwalks forward in a raincoat
Warmer lightsoft golden hour light from the leftsun slowly rising
Match a styleclean studio look, matte skincamera pans up
Keep the facesame face and hair as the stillsmiles and waves

Why actions backfire

The model receives the motion from the clip. A prompt that says the character raises both arms asks for a different motion than the reference shows, and the result is a blend that satisfies neither. At best the words are ignored. At worst they distort the pose.

Camera movement belongs to the reference too. If you want a different camera, change the reference, or set character_orientation to image so that the framing of the still is kept, as the camera and orientation page explains.

A body to start from

Start with the shortest body that works, add a prompt only if the first result needs a fix, and change one thing at a time. A new prompt is a new body, so use a new Idempotency-Key for it, and run it at a short duration_seconds first, since a 1-second probe costs $0.1575 and a 30-second run costs $4.725.

If a face or an outfit keeps drifting, the cause is usually the still, not the prompt. A sharper, front-facing, well-lit image with the whole body in frame works better than any extra words.

A short test plan

Run three probes at 1 second each, for $0.4725 in total. Probe one has no prompt. Probe two adds a one-line appearance prompt. Probe three adds a second line about the light. Compare the three side by side, and keep the shortest prompt that gives the look that you want. Most jobs need no more than one sentence.

Related posts

More in Developers

All Developers posts

Written by Sume