AI video with multiple characters: get everyone in one shot
Give the model every character: one still with the whole cast as the first frame, or one reference image per character on a model that takes several.

To make an AI video with multiple characters, give the video model every character before it starts. Either put the whole cast in one still and use it as the clip's first frame, or send each character's picture as a separate reference image to a model that accepts several. Neither route guarantees that everyone stays recognizable for the whole clip, so watch it before you use it.
Facts come from Sume's Video generation, Video Router, and Image API docs, read on 2026-09-28. Reference caps are the API's current checks. Without code, describe the scene in the Agents tab: the agent picks the models and asks before it spends.
How do I put several characters in one AI video?
There are two documented inputs on POST /v1/videos, and they suit different jobs:
- One still as the first frame. Combine the characters into one image with an image model that takes several references (ChatGPT Image 2.5 takes up to 16), check it, then send it in
frame_imageswithframe_type: "first_frame". The opening frame is then the group you approved. How to combine two photos into one with AI covers the still. - One reference per character. Send each picture in
input_referencesas atype: "image_url"entry and describe the scene in the prompt. The model uses references as visual guidance rather than exact frames. - Not both at once. If a request has
frame_images, it takes precedence and the request is treated as image-to-video; in current code the references are then not sent to the model. - Public HTTPS only. Signed or private URLs are refused as generation inputs, and the Image API's result URLs are signed, so host the still you keep at your own public HTTPS URL first.
Which video models take several character images?
Reference support differs by model. Every model below takes a first frame, so the combined still is an option on all of them.
| Model | Image references | Name each one in the prompt? |
|---|---|---|
gemini-omni-flash-1.1 | Up to 10 | Yes: <IMAGE_REF_0>, <IMAGE_REF_1>, … in list order |
wan-3.0 | Up to 10 | Not documented |
minimax-h3, minimax-h3-max | Up to 9 (12 references of all types) | Not documented |
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-mini | Yes; no documented cap, and the API refuses over 12 references in total | Not documented |
kling-3, grok-imagine-video-1.5 | None | Use one combined still as the first frame |
How do I tell the model which character is which?
On Gemini Omni Flash 1.1, point at each reference by position: <IMAGE_REF_0> is the first image in input_references, <IMAGE_REF_1> the second, counted from 0 in list order. The docs describe these tokens for that model only. On other models, describe each character in words, by clothing or by where they stand, and give each one a clear action.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: two-characters-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "<IMAGE_REF_0> and <IMAGE_REF_1> sit across a cafe table. <IMAGE_REF_0> slides a cup of coffee over and <IMAGE_REF_1> laughs. Static medium shot, warm window light",
"input_references": [
{ "type": "image_url", "image_url": { "url": "https://example.com/character-a.png" } },
{ "type": "image_url", "image_url": { "url": "https://example.com/character-b.png" } }
],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 8
}'Can the characters talk to each other?
Not by laying voices under a generated clip. Sume's docs say video models do not lip-sync to generated speech (TTS) or to a later voice-over. A speaking shot is a lip-sync clip instead: VEED Fabric 1.0 (veed/fabric-1.0), for example, turns one still and an audio track into a talking clip. For a dialogue, make one talking clip per turn and cut them together in order, as AI avatar conversation video: two speakers shows with avatar videos.
Will every character stay recognizable?
Not guaranteed. A first frame pins only the opening picture, everything after it is generated, and references only guide the model. Watch the whole clip and generate again if a character changes or drops out. To keep the same characters across several shots, see Consistent character across AI video shots.
What are the limits?
- Reference caps per model are in the table;
kling-3andgrok-imagine-video-1.5take none. - Clip length depends on the model: 3–10 seconds on
gemini-omni-flash-1.1, up to 30 onseedance-2.5andwan-3.0, and at most 15 on the other models. - Every image URL must be public HTTPS.
- Clips are billed per model at provider list × 1.25, reserved on submit; each model's rates are in
pricing_skusonGET /v1/videos/models.
Sources
Related posts
More in Models
- How to change the color of an object in an image with AI
Send the photo to an AI image model, name the object and its new color, and list what must stay. One request per color, checked against the original.
- How to change the background of a video with AI
Give a video-to-video AI model your clip and a prompt that names the new background and what must stay. How to do it on Sume and what to check.
- Day to night video with AI: edit, generate, or grade
Turn a day video into night with a video-to-video AI edit, generate a day-to-night clip from two stills, or darken it with a color grade.
- Dreamina vs Seedance: the app and the model it runs
Dreamina vs Seedance: Seedance is ByteDance Seed's video model; Dreamina is a creative app that runs it. When to use the app and when an API.
Written by Sume