Kling 4.0 stable motion vs Sume's Kling 3.0 Motion Control
Kling 4.0 claims stable motion for dance and sports. To copy a specific move on Sume, use Kling 3.0 Motion Control with a driving video of up to 30 s.

Kling 4.0's page says it has enhanced motion stability for action, sports and dance (read 2026-10-02), but that describes text and reference generation, not copying a specific move. To copy a given movement onto a still on Sume, call Kling 3.0 Motion Control (kling/3.0/motion-control) with a driving video of up to 30 seconds. Sume lists no Kling 4.0 id.
What does Kling 4.0 say about motion?
The Kling 4.0 feature page lists stable dynamic motion for action sequences, sports, dance and extreme sports, 30-second single-pass clips, up to ten keyframes, and up to 15 combined references. Flash is limited to 20 seconds and 720p, and is in a limited rollout.
Those are properties of a generative model steered by prompts and references. The page does not say you can hand it a video and have that motion transferred onto a photo.
| Need | Kling 4.0 page | Sume Motion Control |
|---|---|---|
| Copy a move onto a still | Not stated | Yes: still plus driving video |
| Driving or clip length | Up to 30 s clip | Driving video up to 30 s |
| Model id | Kling 4.0, rolling out | kling/3.0/motion-control |
| Prompt role | Steers the generation | Steers appearance only |
How does Sume's Motion Control work?
You give a public HTTPS still (or a ready avatar), a public HTTPS motion video, and duration_seconds. The output length follows the driving video. character_orientation decides whose framing wins: video (default) or image. keep_original_sound keeps the driving audio by default.
The optional prompt is up to 2000 characters and, per the schema, only steers appearance details; motion and framing follow the video.
Which should I choose?
If you can film or find the exact move, Motion Control is the controlled route, and it exists on Sume today. If you want a model to invent the motion from a prompt, use the video models Sume lists and check the video catalog for ids.
Read Jobs and results for the poll-and-fetch loop that every job uses.
Sources
Related posts
More in Comparisons
- Kling @Element character binding vs Sume: what to use instead
Kling v3 Pro on fal binds characters with elements. Sume's kling-3 takes frames only, so use a reference-capable model or one fixed first frame.
- Kling v3 Pro multi_prompt shots on fal vs Sume's one clip per shot
fal's Kling v3 Pro has a multi_prompt list for multi-shot video. Sume's kling-3 sends one prompt and frames per clip, so you join shots with Timeline 1.0.
- Kling v3 Turbo Pro at $0.14 a second vs v3 Pro and Sume's kling-3
fal lists Kling v3 Turbo Pro at $0.14 a second and v3 Pro at $0.112 audio off, $0.168 on. What that means for 5 seconds, and the Kling id Sume lists.
- Leonardo.Ai API alternative for image generation: Sume Images
Leonardo's API submits to /generations and returns a generationId. Sume POST /v1/images works from a model catalog with idempotent retries. Compared.
Written by Sume