AI video multiple faces consistency: Luma's 8-face limit, Sume's caps
Luma Ray3.2 lists facial performance tracking for up to 8 faces. Sume has no such setting; the limit is how many reference images each video model accepts.
Luma's Ray3.2 announcement lists facial performance tracking for up to 8 faces in one clip. Sume's video API has no face-tracking setting, so there is no face count to hit; the number that limits a many-face shot is how many reference images the chosen model accepts, from 9 or 10 images up to 12 files in total.
The Luma figure is from its Ray3.2 post, dated June 9 2026. The Sume figures are from the video docs and the API's own request checks, read 2026-09-29.
What does Luma's 8-face limit describe?
It describes Luma's own tracking feature in Ray3.2, which the post says is available as an API. It says nothing about other models, and this page does not compare quality with it. If your shot has more than 8 people, Luma's stated cap is where that feature ends.
What limits faces on Sume instead?
On Sume you give a model a still for each person through input_references, and the model uses them as visual guidance rather than exact frames. So the ceiling is the reference cap of the model you pin.
| Model id | Reference images | Total cap |
|---|---|---|
seedance-2.5 | Counted with videos and audio | 12 files |
wan-3.0 | Up to 10 | Counted per type |
minimax-h3, minimax-h3-max | Up to 9 | 12 files |
gemini-omni-flash-1.1 | Up to 10 | Counted per type |
kling-3 | No reference inputs | Not applicable |
How do I keep many faces consistent?
Use one clear, front-lit still per person, and describe who stands where in the prompt. A caps table only tells you the maximum; a crowd of more people than the model takes needs a different plan: split the scene into shots with fewer people each and join the clips.
Two routes exist for a large group. Compose the group into one still and send it as first_frame, or generate each shot with its own small set of references. Video with multiple characters and character consistency across shots cover both.
What if the scene has more people than the cap?
Plan the scene as a sequence. A wide shot can show the whole group at a distance, where individual faces matter less, while close-ups use two or three references each. Generate the shots as separate jobs, keep the same reference stills for the same person in every shot, and join the clips afterwards. This also keeps each request under the model's cap and makes a bad shot cheap to redo, because you rerun one clip rather than the whole scene.
Does more references mean better faces?
Not necessarily. The caps are ceilings, not targets. Test with the fewest references that carry the look, add one at a time, and check each result before you scale up the number of people.
Sources
Related posts
More in Models
- Why AI video models do not lip sync to a voiceover, and what to use
Sume docs: video models do not lip-sync to generated speech or a voice-over. For a talking face use Avatar video or the lip-sync endpoint (still plus audio).
- Batch edit photos with AI: one edit across many photos
To batch edit photos with AI, send the same edit prompt once per photo, with that photo as the reference. How to script it, pace it, and what it costs.
- Cartesia Sonic 3.6 API: which model id to send
Send sonic-3.6 for the latest stable Sonic 3.6, a dated id to freeze it, or sonic-preview for beta. Which of these ids Sume's TTS Router lists.
- AI outfit change video: change clothes in a clip with AI
Change someone's outfit in an existing video with an AI video-to-video edit: name the garment, what it becomes, and what must stay the same.
Written by Sume