Does Kling motion control follow my prompt? Only for appearance
On Sume's Kling 3.0 Motion Control the optional prompt (up to 2,000 characters) steers appearance only; motion and framing follow the driving video.

No: on Sume's Kling 3.0 Motion Control, the optional prompt does not direct the movement. The API reference (read 2026-10-09) says that motion and framing follow the driving video and that the prompt only steers appearance details. The field accepts 1 to 2,000 characters. If you want the avatar to move differently, change the driving video, not the prompt.
What each input controls
Motion control separates three jobs. The still decides who appears. The driving video decides how they move and how long the clip runs. The prompt nudges how the result looks. Two flags handle the rest.
| Field | Controls | Limits |
|---|---|---|
| image_url or avatar_id / avatar_handle | Who appears | Send exactly one |
| motion_video_url | Motion, framing, output length | Public HTTPS, max 30 s |
| prompt | Appearance details only | 1 to 2,000 characters, optional |
| character_orientation | Whose framing wins | video (default) or image |
| keep_original_sound | Driving audio in output | true by default |
| duration_seconds | Credit reservation | 1 to 30 |
What a useful prompt looks like
Because the model already has the motion, spend prompt words on what the video cannot tell it: the outfit's color, the lighting you want to match, the material of a prop. Keep it short and concrete. A prompt such as 'matte red jacket, soft studio light' does the job, and a prompt like 'spin twice and wave' asks for motion that the driving video has to supply.
Prompts do not change the price. Cost is $0.1575 per output second on the catalog, billed on ceil of the motion-video second, so a 14-second reference is 14 x 0.1575 = $2.205 whatever the prompt says.
If the motion is wrong
Record or pick a new driving video. Test a short 3-second piece first: 3 x 0.1575 = $0.4725 shows whether the movement transfers to your still before you pay $4.73 for 30 seconds. Choose character_orientation: image when you want the still's framing to hold, and video when the dancer's framing should win. If the driving clip has speech or music you cannot reuse, set keep_original_sound to false.
The prompt is also not a safety valve. If the output is not what you want, rewording it cannot fix a mismatch between the still and the reference movement.
Testing a prompt cheaply
Because prompts do not change motion, you can test a prompt on a short driving clip. Cut 3 seconds from your reference ($0.4725), try two prompts, compare the look, and only then run the full 14-second clip at $2.205. That is a $0.945 test for two prompt variants against the risk of paying $4.41 for two wrong full-length renders.
If you use an avatar instead of a still, remember it resolves to the avatar's identity still, so appearance prompts may have less to steer. Use an image_url when you need to vary the look per run.
Sources
Related posts
More in Media tools
- Korean captions: slam returns 400, so use korean-ad with language ko
Latin caption styles have no Hangul glyphs, so Korean text on slam, punch or tiktok-green returns 400 caption_hangul_text_latin_style. Use korean-ad.
- LinkedIn 1200x675 video ad: Sume needs even sizes, so use 1280x720
LinkedIn lists 1200x675 as a 16:9 size, but 675 is odd. Sume timeline sizes must be even, so render 1280x720 or 1920x1080, both inside LinkedIn's range.
- LinkedIn 30-minute video: 24 spot-check stills, one every 75 seconds
LinkedIn allows videos up to 30 minutes. Video inspect takes 1,800 seconds and 24 stills per call, so at[] with 37.5 + 75k covers the whole file in one job.
- LinkedIn 30-minute video ad at 500 MB: 2.2 Mbps average, $3.00 render
A 30-minute LinkedIn ad must average under about 2.2 Mbps to fit 500 MB. Sume's timeline renders up to 1,800 seconds for $3.00; trim stops at 900.
Written by Sume