Luma Ray3.2 face tracking vs Kling motion control: pick by input
Luma Ray3.2 tracks up to eight faces frame by frame; Kling motion control moves a still with a reference clip. Which to use, and what Sume runs today.

Use Luma Ray3.2 when you want a generated shot that keeps an actor's expressive performance, and use Kling 3.0 Motion Control when you already have a still and a clip whose movement it should copy. They start from different inputs. On Sume you can run the Kling route today; Ray3.2 is not in Sume's video catalog.
Luma released Ray3.2 on June 9, 2026, and says its API is available for the first time with this model (read 2026-10-07).
What Luma says Ray3.2 does
Luma's announcement lists the control surface below. It does not give prices or credit costs on that page, so none are quoted here.
| Capability | Luma states | Closest Sume option |
|---|---|---|
| Keyframes | Up to 16 keyframes in a single clip | First and last frame on models that list them |
| Length and size | Clips up to 20 seconds at 1080p | Seedance 2.5 and Wan 3.0 go to 30 seconds |
| Output formats | Native HDR and 16-bit EXR export | MP4 artifacts |
| Faces | Full expressive state for up to eight faces, frame by frame | No equivalent field |
| API | Available as an API | Not in the Sume catalog |
What Kling motion control does instead
Kling's route does not generate a scene from a prompt. It takes a still of a character and a motion reference video, then animates the still with the movement in the clip. The output length follows the reference clip, from 3 to 30 seconds by Kling's guide (read 2026-10-07).
On Sume, POST /v1/kling/3.0/motion-control takes image_url or a ready avatar, a public HTTPS motion_video_url, and duration_seconds. It bills $0.1575 per second, and it keeps the reference clip's audio by default (keep_original_sound).
Pick by the input you already have
- You have only a script and want a cinematic shot with controlled keyframes: Ray3.2 is built for that, and Sume cannot run it.
- You have a still of a character and a phone clip of the move you want: Kling motion control.
- You need to put a different person into an existing video, keeping its motion and sound: H3 Max Recast on Sume takes one to four photos.
- You need a talking still driven by audio: that is a lip-sync route, not either of these.
What a Sume run looks like
A 10-second reference reserves $1.575, captured on completion and refunded on failure. Output is an MP4, with no HDR or EXR delivery on this route.
If Ray3.2's 16 keyframes are the reason you came, the nearest Sume approach is chained jobs with first and last frames, which is described in 16 keyframes against chained jobs. It is a workaround, not a match.
Sources
Related posts
More in Comparisons
- 30-second lyric or music clip with an audio reference: $8.07 vs $1.88
Both seedance-2.5 and wan-3.0 take an audio reference on Sume. A 30-second clip is $8.07 at 480p on Seedance 2.5 and $1.88 at 480p on Wan 3.0.
- MAI-Transcribe-2-Streaming partials vs Sume STT terminal webhook
MAI-Transcribe-2-Streaming sends intermediate and final results as audio streams in. Sume STT sends one signed terminal callback. How to design for each.
- MAI custom voice SSML (ttsembedding) vs Sume voice ids: what is gated
Microsoft's MAI-Voice-2.1 custom voice sits behind Limited Access Review and uses a speakerProfileId in SSML. Sume picks voices by id, with no clone widget.
- Edit models Morphic names versus what the Sume catalog lists
Morphic runs image edits on Seedream 5.0 Pro, Nano Banana 2 and GPT Image 2.5 until Ideogram 4.5 arrives. Which of these Sume lists, and where to check.
Written by Sume