Swap multiple characters in a video with AI: 2 to 4 people
Swap several people in one clip with H3 Max Recast on Sume: one photo per person, left to right by default, up to four, and a price that ignores headcount.

To swap more than one person in a video, send H3 Max Recast one reference photo per new person, in the order they should land. Sume's h3-max-recast accepts one source clip of 5 to 30 seconds and 1 to 4 photos in reference_image_urls. fal's schema says the photos replace the main people in the video from left to right, each in every shot they appear in.
So a two-person interview gets two photos: the photo of whoever should replace the person on the left first, the other second. Headcount does not change the bill. Sume prices the job by source seconds and resolution only, and fal says reference images are included at no additional cost.
How does Recast decide who becomes whom?
By default it is positional. fal's schema for reference_image_urls says the photos replace the main people in the video from left to right. Sume passes your list through in the order you send it, so the first URL is the leftmost main person.
Position can be the wrong key. Two people can swap sides mid-shot, or the person who matters can stand on the right. For that, send a prompt. Sume allows one of up to 2000 characters, described in its OpenAPI as who becomes whom or what else to keep or change, and fal calls it optional guidance on character mapping. Neither page gives a prompt grammar, so treat the wording as something to test on a short clip first.
What does a two-person request look like?
This is a Video Router call. duration is not the length you want back: Sume requires the inspected length of your source, in whole seconds rounded up, because the output keeps the source length. Read it with ffprobe first (a 11.2 second clip is 12).
Every URL must be public HTTPS. Signed, private and localhost URLs are rejected before any paid work starts.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recast-duo-001" \
-d '{
"model": "h3-max-recast",
"video_url": "https://example.com/duo.mp4",
"reference_image_urls": [
"https://example.com/left-person.png",
"https://example.com/right-person.png"
],
"prompt": "The person on the left becomes the first photo; the person on the right becomes the second.",
"resolution": "768p",
"duration": 12,
"mode": "async"
}'What are the hard limits for a multi-person swap?
Limits from the Sume catalog entry and fal's schema, read 2026-10-03:
- Sume's catalog says Recast replaces the main people. Neither page promises that bystanders are swapped, so check crowds in the output.
- A photo that does not match a person on screen is a mapping you have to describe in the prompt.
aspect_ratio,generate_audio, frame images and audio references are all rejected; see the full rejected-field list.
| Limit | Value | Where it comes from |
|---|---|---|
| Photos | 1 to 4, one per new person | Sume catalog; five or more are rejected |
| Source length | 5 to 30 seconds | Sume catalog and fal schema |
| Longest single shot | 15 seconds | fal schema |
| Resolution | 768p (Sume default) or 1080p | Sume catalog; fal's own default is 1080P |
| Prompt | Optional, up to 2000 characters | Sume OpenAPI |
| Who is swapped | The main people in the video | Sume catalog and fal schema |
| Price | Per source second, not per person | Sume billing: list x 1.25 |
How should I prepare the photos and the clip?
Both vendors are thin on photo guidance: fal says one photo per new person and nothing about framing, and Sume's docs add only the URL rules. What follows is our practical suggestion, not a vendor requirement. Give each person one clear photo of that person alone, so the mapping is unambiguous, and keep the file at a stable public HTTPS address until the job finishes, because the job has to fetch it when it runs.
For the source, pick a clip where the main people stay visible. A clip that is mostly one wide shot of a crowd gives the model little to map. If your footage runs past 30 seconds or has one very long take, read the 30 second trim recipe first, because fal also forbids any single shot over 15 seconds.
What does a multi-person clip cost?
Sume bills fal's per-second list times 1.25: $0.375 a second at 768p and $0.5625 at 1080p. The 12 second duo above is $4.50 at 768p (12 x $0.375) and $6.75 at 1080p. A solo 12 second clip costs exactly the same, so swapping four people is not four times the price.
The cost that does multiply is retries. Because the model is not seedable on Sume, plan on more than one pass whenever the mapping matters, and test the prompt on the shortest valid slice (5 seconds, $1.88 at 768p) before you spend on the full clip.
What should I check in the result?
Watch the first appearance of each person, then any shot where two people cross. fal says each photo replaces its person in every shot they appear in, which is the part worth confirming frame by frame. Source audio is kept, so voices still belong to the original speakers; see whether Recast keeps the audio.
Sume has no seed on this model, so a clip where two people got swapped the wrong way round cannot be re-rolled identically. Resubmit with a clearer prompt and a new Idempotency-Key. Each submit is a new paid job at the same per-second price.
If you only need to change one face and not the person, Sume's Beta face swap is a different tool with its own limits. Only submit photos of people who agreed to appear. Sume's Terms of Service make you responsible for permission to use any person's likeness in content you submit.
Sources
Related posts
More in Use cases
- Template bulk edits on YouTube Shorts: what to vary per row
YouTube's Oct 1, 2026 originality update names template-based bulk changes as not original. How to make each row of a Sume bulk run differ in substance.
- Where GMV Max creative videos come from, and where AI clips go
TikTok Product GMV Max only runs videos from connected accounts, Spark Ads posts or ACA, with a product anchor link. Where a Sume-made MP4 has to go first.
- If TikTok removes an AI clip, are edited variants caught too?
TikTok says it will try to catch similar versions of misleading AI video that violates policy. What that scope is, and why bulk variants are a risk.
- UGC ad casting test: one script across three AI avatars
Test which AI spokesperson sells your script: create three Sume previews with the same copy, pick a still, then render only the winner. Curl loop and costs.
Written by Sume