AI video with multiple characters: get everyone in one shot

Give the model every character: one still with the whole cast as the first frame, or one reference image per character on a model that takes several.

5 min readSume
All posts

To make an AI video with multiple characters, give the video model every character before it starts. Either put the whole cast in one still and use it as the clip's first frame, or send each character's picture as a separate reference image to a model that accepts several. Neither route guarantees that everyone stays recognizable for the whole clip, so watch it before you use it.

Facts come from Sume's Video generation, Video Router, and Image API docs, read on 2026-09-28. Reference caps are the API's current checks. Without code, describe the scene in the Agents tab: the agent picks the models and asks before it spends.

How do I put several characters in one AI video?

There are two documented inputs on POST /v1/videos, and they suit different jobs:

  • One still as the first frame. Combine the characters into one image with an image model that takes several references (ChatGPT Image 2.5 takes up to 16), check it, then send it in frame_images with frame_type: "first_frame". The opening frame is then the group you approved. How to combine two photos into one with AI covers the still.
  • One reference per character. Send each picture in input_references as a type: "image_url" entry and describe the scene in the prompt. The model uses references as visual guidance rather than exact frames.
  • Not both at once. If a request has frame_images, it takes precedence and the request is treated as image-to-video; in current code the references are then not sent to the model.
  • Public HTTPS only. Signed or private URLs are refused as generation inputs, and the Image API's result URLs are signed, so host the still you keep at your own public HTTPS URL first.

Which video models take several character images?

Reference support differs by model. Every model below takes a first frame, so the combined still is an option on all of them.

From Video generation and Video Router, read 2026-09-28. Caps are the API's current checks; GET /v1/videos/models lists the reference types each model takes, not the caps.
ModelImage referencesName each one in the prompt?
gemini-omni-flash-1.1Up to 10Yes: <IMAGE_REF_0>, <IMAGE_REF_1>, … in list order
wan-3.0Up to 10Not documented
minimax-h3, minimax-h3-maxUp to 9 (12 references of all types)Not documented
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-miniYes; no documented cap, and the API refuses over 12 references in totalNot documented
kling-3, grok-imagine-video-1.5NoneUse one combined still as the first frame

How do I tell the model which character is which?

On Gemini Omni Flash 1.1, point at each reference by position: <IMAGE_REF_0> is the first image in input_references, <IMAGE_REF_1> the second, counted from 0 in list order. The docs describe these tokens for that model only. On other models, describe each character in words, by clothing or by where they stand, and give each one a clear action.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: two-characters-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "<IMAGE_REF_0> and <IMAGE_REF_1> sit across a cafe table. <IMAGE_REF_0> slides a cup of coffee over and <IMAGE_REF_1> laughs. Static medium shot, warm window light",
    "input_references": [
      { "type": "image_url", "image_url": { "url": "https://example.com/character-a.png" } },
      { "type": "image_url", "image_url": { "url": "https://example.com/character-b.png" } }
    ],
    "aspect_ratio": "16:9",
    "resolution": "720p",
    "duration": 8
  }'

Can the characters talk to each other?

Not by laying voices under a generated clip. Sume's docs say video models do not lip-sync to generated speech (TTS) or to a later voice-over. A speaking shot is a lip-sync clip instead: VEED Fabric 1.0 (veed/fabric-1.0), for example, turns one still and an audio track into a talking clip. For a dialogue, make one talking clip per turn and cut them together in order, as AI avatar conversation video: two speakers shows with avatar videos.

Will every character stay recognizable?

Not guaranteed. A first frame pins only the opening picture, everything after it is generated, and references only guide the model. Watch the whole clip and generate again if a character changes or drops out. To keep the same characters across several shots, see Consistent character across AI video shots.

What are the limits?

  • Reference caps per model are in the table; kling-3 and grok-imagine-video-1.5 take none.
  • Clip length depends on the model: 3–10 seconds on gemini-omni-flash-1.1, up to 30 on seedance-2.5 and wan-3.0, and at most 15 on the other models.
  • Every image URL must be public HTTPS.
  • Clips are billed per model at provider list × 1.25, reserved on submit; each model's rates are in pricing_skus on GET /v1/videos/models.

Sources

Related posts

More in Models

All Models posts

Written by Sume