AI video multiple faces consistency: Luma's 8-face limit, Sume's caps

Luma Ray3.2 lists facial performance tracking for up to 8 faces. Sume has no such setting; the limit is how many reference images each video model accepts.

4 min readSume
All posts

Luma's Ray3.2 announcement lists facial performance tracking for up to 8 faces in one clip. Sume's video API has no face-tracking setting, so there is no face count to hit; the number that limits a many-face shot is how many reference images the chosen model accepts, from 9 or 10 images up to 12 files in total.

The Luma figure is from its Ray3.2 post, dated June 9 2026. The Sume figures are from the video docs and the API's own request checks, read 2026-09-29.

What does Luma's 8-face limit describe?

It describes Luma's own tracking feature in Ray3.2, which the post says is available as an API. It says nothing about other models, and this page does not compare quality with it. If your shot has more than 8 people, Luma's stated cap is where that feature ends.

What limits faces on Sume instead?

On Sume you give a model a still for each person through input_references, and the model uses them as visual guidance rather than exact frames. So the ceiling is the reference cap of the model you pin.

Reference-image caps per model from Sume's API checks and catalog, read 2026-09-29. Confirm with GET /v1/videos/models before you submit.
Model idReference imagesTotal cap
seedance-2.5Counted with videos and audio12 files
wan-3.0Up to 10Counted per type
minimax-h3, minimax-h3-maxUp to 912 files
gemini-omni-flash-1.1Up to 10Counted per type
kling-3No reference inputsNot applicable

How do I keep many faces consistent?

Use one clear, front-lit still per person, and describe who stands where in the prompt. A caps table only tells you the maximum; a crowd of more people than the model takes needs a different plan: split the scene into shots with fewer people each and join the clips.

Two routes exist for a large group. Compose the group into one still and send it as first_frame, or generate each shot with its own small set of references. Video with multiple characters and character consistency across shots cover both.

What if the scene has more people than the cap?

Plan the scene as a sequence. A wide shot can show the whole group at a distance, where individual faces matter less, while close-ups use two or three references each. Generate the shots as separate jobs, keep the same reference stills for the same person in every shot, and join the clips afterwards. This also keeps each request under the model's cap and makes a bad shot cheap to redo, because you rerun one clip rather than the whole scene.

Does more references mean better faces?

Not necessarily. The caps are ceilings, not targets. Test with the fewest references that carry the look, add one at a time, and check each result before you scale up the number of people.

Sources

Related posts

More in Models

All Models posts

Written by Sume