Kling Motion Control prompt: appearance only, motion is the video
On Sume's Kling 3.0 Motion Control the prompt guides appearance only. Motion comes from motion_video_url, and the body rejects reference_image_urls and model.

In Sume's Kling 3.0 Motion Control, the prompt steers how the character looks. It does not choose the movement: motion_video_url drives all output motion. If the clip moves wrong, change the motion video, not the prompt.
What each field controls
Each field has one job in the Sume contract for POST /v1/kling/3.0/motion-control.
| Field | Job |
|---|---|
motion_video_url | Required. Drives all output motion |
image_url or avatar_id/avatar_handle | Exactly one; the character still |
prompt | Optional. Appearance steering only |
duration_seconds | 1-30. Reservation basis only; never forwarded |
keep_original_sound | Default true; false returns a silent clip |
character_orientation | video (default) or image |
Duration and price
The provider has no duration setting: the length of the motion video sets the output length. Sume only uses duration_seconds to price the job at admission, as ceil(duration_seconds) x $0.1575 (fal list $0.126 x 1.25). If you declare 10 s but upload a 6 s motion video, the clip is 6 s long while the hold was for 10.
What the vendor schema says
fal's own schema calls the prompt an optional description of the action or scene, with an example of a man dancing (fal page, read 2026-10-05). Sume's doc is stricter: it positions the prompt as appearance guidance. Write your prompt to describe the person and style, and let the reference video carry the action.
Strict body
This route is strict. It rejects Video 1.0 vocabulary such as reference_image_urls, reference_video_urls, model and provider endpoint fields. That keeps it from becoming a back door into the kling-3 text-to-video reference inputs, which do not exist on that row either.
Example
A body that sets appearance in the prompt and leaves motion to the video.
{
"image_url": "https://example.com/character.png",
"motion_video_url": "https://example.com/walk-6s.mp4",
"duration_seconds": 6,
"prompt": "Same person, navy raincoat, soft overcast light"
}Sources
Related posts
More in Media tools
- Lip-sync a saved avatar with H3 Max: avatar_handle, not image_url
Use a ready Sume avatar as the face for MiniMax H3 Max Lip Sync: send avatar_handle plus Sume-hosted audio, 5-14.8 seconds, and see the price at 768p.
- Put a logo or lower third over a talking avatar video with compose
Timeline compose overlay places a Sume-hosted still over an avatar clip, at the top, center or bottom. $0.02 flat per job. The layout keys and a request.
- Use a MAI-Voice-2.1 clip as avatar audio? Sume needs its own file
A talking still on Sume takes Sume-hosted audio under 10 MB. A MAI-Voice-2.1 clip sits elsewhere, so make the line with Sume TTS instead. Sizes and limits.
- Match voice emotion and music mood: one mood word, two Sume fields
Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.
Written by Sume