Replace an actor in a video with AI: what Recast keeps and swaps
To replace an actor in a video with AI on Sume, send the clip and one photo per person to h3-max-recast. Motion, camera, cuts and sound stay; people change.

To replace an actor in a finished video with AI, use the h3-max-recast model on Sume: you send the source clip as video_url and one photo per new person as reference_image_urls (1 to 4), and the output keeps the source motion, camera, cuts and sound while the people on screen change (Video Router docs, read 2026-10-03). Nothing is re-shot and no text prompt is required, because the prompt is optional.
The model is MiniMax H3 Max running as fal Recast. fal describes it as changing who is on screen while preserving the original motion, camera movements, cuts and audio (read 2026-10-03). Sume lists it in the Video Router catalog, so you call it with the same API key and wallet as every other video model, billed at the provider list price times 1.25.
What stays and what changes
Think of Recast as a person-for-person substitution on a clip you already like. The performance, the framing and the edit are the source; the identity is the variable. That makes it the right tool when a take is good but the casting is wrong, and the wrong tool when you want a different performance, a different location or a different shot list. For those, generate a new clip instead.
| Aspect | Documented behavior | Source |
|---|---|---|
| People on screen | Replaced by the people in your reference photos, one photo per person | Sume Video Router docs; fal API schema |
| Motion and camera | Kept from the source video | fal model page |
| Cuts | Kept; no single shot in the source may exceed 15 seconds | fal API schema; Sume Video Router docs |
| Audio | Source audio is kept in the output | fal API schema |
| Source length | 5 to 30 seconds; output duration is the source length | Sume Video Router docs |
| Photos | 1 to 4 images, one per person | Sume Video Router docs |
| Resolution | 768p (Sume default) or 1080p | Sume Video Router docs |
| Prompt | Optional, up to 2,000 characters | Sume docs/api video-router table |
A request that replaces one actor
Send the source and one photo of the replacement. With one person on screen and one photo you do not need a prompt. The call below is the documented Video Router shape; Idempotency-Key makes a retry return the original job instead of a second charge.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recast-actor-001" \
-d '{
"model": "h3-max-recast",
"video_url": "https://example.com/scene.mp4",
"reference_image_urls": ["https://example.com/new-actor.jpg"],
"resolution": "768p",
"mode": "async"
}'What you pay and what can stop the job
Billing is per second of source video. fal lists Recast at $0.30 per output second at 768p and $0.45 at 1080p, with reference images included at no extra charge (read 2026-10-03). Sume's catalog applies a 1.25 multiplier, so a 10-second source costs about $3.75 at 768p and about $5.63 at 1080p before any rounding. For the per-clip arithmetic see what a 12-second clip bills.
Three limits decide whether a clip can be recast at all. The source must be 5 to 30 seconds. No single shot may run longer than 15 seconds, so a locked-off 20-second take is out until you split it. And the model is source-video-only: Sume rejects a text-only request to h3-max-recast, and sume/auto never routes to it, so you must name the id yourself.
- Public HTTPS URLs only for the video and the photos.
- One photo per person; photos map to people left to right by default, and the prompt can override that.
- No
aspect_ratiofield: output follows the source.
Before you submit a replacement
A short checklist saves a failed job and a retry. First, probe the source: confirm it is between 5 and 30 seconds and look at where the cuts fall, because a shot longer than 15 seconds will not be accepted. The video inspect docs describe a probe plus stills call you can run on a Sume-hosted clip. Second, count the people you want replaced and gather exactly that many photos; the default mapping is left to right across the frame, so order your list the way the cast stands, or say who becomes whom in the prompt. Third, choose a resolution: 768p is the default and the cheaper tier, and a draft at 768p is a sensible way to check the casting before paying for 1080p.
Finally, decide what happens to the voice. The output keeps the source audio, so a replaced actor will still sound like the original performer. That is a feature when the line reading is the thing you love and a problem when the new person should sound different. Sume does not change that in the Recast call; plan the voice as a separate step.
When to pick something else
If the new person should be a face on your own saved avatar and the clip is 4 to 15 seconds with usable audio, the Beta face swap endpoint applies one ready avatar face. If you want a photo to perform a motion from a reference video rather than replace someone in footage, that is Kling motion control. Comparing these is covered in face swap vs Recast. And for a replaced person who must not be a real individual who has not agreed, do not use another person's likeness without permission; Sume's docs do not make a legal determination for you.
Sources
Related posts
More in Use cases
- Replace someone with yourself in a video with AI: Recast on Sume
Put yourself into an existing clip: send a public source video and one photo of you to H3 Max Recast on Sume. Limits, cost per clip and consent rules.
- Reply to a YouTube comment with a Short: the comment sticker
YouTube lets you answer a comment with a Short and shows the comment as a sticker. How to start one, and how Sume prepares the vertical reply clip.
- Restyle a video as claymation or watercolor with Gemini Omni
Apply a claymation or watercolor look to existing footage with Gemini Omni Flash 1.1 on Sume: Google's style wording, the edit request, and a still check.
- Restyle ad captions without transcribing the video twice
Pass source_caption_id to Sume's video-captions endpoint to re-burn an ad in a new style, reusing the first run's word timings with no second speech-to-text.
Written by Sume