Ray 3.2 tracks 8 faces; Sume swaps 1-4 people with H3 Max Recast
Luma Ray 3.2 lists facial tracking for up to 8 faces. Sume has no tracking output, but h3-max-recast swaps 1-4 people in a clip and Kling drives a still.

Luma lists facial performance tracking for up to 8 faces in Ray 3.2. Sume does not run Ray 3.2 and returns no face-tracking data, but it has two face-related tools: H3 Max Recast replaces 1 to 4 people in a source clip with people from photos, and Kling 3.0 Motion Control drives a still image with the motion of a reference video.
Those do different jobs. Tracking is about following performances in a generation; recast and motion control are about changing who appears in a shot.
What Luma lists
The Ray 3.2 launch page (read 2026-10-05) lists facial performance tracking, up to 8 faces, alongside 16 keyframes, 1080p and clips up to 20 seconds. The page is the only source used here for Ray 3.2 claims.
Recast: replace the people in a clip
In the Video Router docs, h3-max-recast replaces the people in a source video_url with the people from 1 to 4 reference_image_urls, with one photo per person. It runs at 768p or 1080p, accepts 5 to 30 seconds, and takes an optional prompt. The duration is the source length.
If your source has two people, send two photos. The model replaces each person, which is closer to a casting swap than to tracking.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recast-001" \
-d '{"model":"h3-max-recast",
"video_url":"https://example.com/source.mp4",
"reference_image_urls":["https://example.com/person-a.jpg","https://example.com/person-b.jpg"],
"resolution":"768p","mode":"async"}'Motion control: a still performs a reference
Kling 3.0 Motion Control (POST /v1/kling/3.0/motion-control, public id kling/3.0/motion-control) takes one still (image_url, or a ready avatar) and a motion_video_url. The reference video drives all motion in the output. The hosted MCP tool is kling-motion-control_create, listed in the MCP tools and gates docs.
This is the nearest Sume answer to a performance transfer, though it is not facial analysis: you get a video, not tracking data.
Side by side
| Capability | Ray 3.2 | Sume |
|---|---|---|
| Face count | up to 8 tracked | recast: 1-4 people, one photo each |
| Output | video | MP4 video; no tracking data |
| Input | Luma generation | a source video plus photos (recast), or a still plus motion video (Kling) |
| Clip length | up to 20 s | recast 5-30 s; Kling follows the motion video, declared 1-30 s |
| Runs on Sume | no | yes |
Choosing between the two Sume paths
Use recast when you already have footage and want different people in it. Use Kling motion control when you have a character still and a performance video, and want the still to move like the performance. If you need more than 4 people, split the shot into separate jobs and join them in Timeline.
Neither tool promises identity preservation beyond what the model does with your photos, so test on a short clip before you commit a long one. Each job reserves its price at admit and uses the Sume list times 1.25 rule, so a short test is cheap.
A cost check before a recast run
H3 Max Recast is billed at the provider list times 1.25 like each other catalog model, and its duration is the source length, so a 12-second source means 12 billed seconds. Trim the source to the shot you need first with video trim to avoid paying for seconds you will cut anyway.
Recast accepts 5 to 30 seconds. A source shorter than 5 seconds is outside the range, so pad or choose a longer range.
What to verify in your own test
- Does each replaced person keep the right position in the frame across the whole clip?
- Does a person who leaves and re-enters the frame keep the same photo identity?
- Does the clip's audio survive, or do you re-add it in Timeline?
- Do hands and profile angles hold up at 768p versus 1080p?
What this page does not claim
It does not claim that Sume tracks faces, that recast matches Ray 3.2's output, or that either tool preserves an identity to a given standard. It reports what the docs say each Sume surface accepts and what Luma's own page lists. Run a short test with your own footage and photos before you commit to a campaign.
Sources
Related posts
More in Comparisons
- Reference limits per request, vendor vs Sume: Wan 3.0, Seedance, Omni
What each vendor says one video request can take, next to what Sume's catalog allows: Wan 3.0, Seedance 2.5, Gemini Omni 1.1 Flash, and Ray3.2.
- Replicate's default webhook secret versus Sume's signing secret
Replicate's secret is at /v1/webhooks/default/secret and it signs id.timestamp.body. Sume's is at /v1/webhooks/signing-secret and signs timestamp.raw_body.
- Retouch 200 product photos: Image API, Agent Completion or Format?
For 200 identical retouches use POST /v1/images: fixed model, price per image, plan queue limits. Use an Agent Completion only if each photo needs judgement.
- Rime Mist v3 vs Coda vs Sume TTS: cost for 1M characters
Rime lists Mist v3 at $0.03 and Coda at $0.05 per 1K characters. Sume TTS 1.0 is $0.0475. Here is the cost of 1M characters and a 20,000-character call.
Written by Sume