Replace an actor in a video with AI: what Recast keeps and swaps

To replace an actor in a video with AI on Sume, send the clip and one photo per person to h3-max-recast. Motion, camera, cuts and sound stay; people change.

6 min readSume
All posts

To replace an actor in a finished video with AI, use the h3-max-recast model on Sume: you send the source clip as video_url and one photo per new person as reference_image_urls (1 to 4), and the output keeps the source motion, camera, cuts and sound while the people on screen change (Video Router docs, read 2026-10-03). Nothing is re-shot and no text prompt is required, because the prompt is optional.

The model is MiniMax H3 Max running as fal Recast. fal describes it as changing who is on screen while preserving the original motion, camera movements, cuts and audio (read 2026-10-03). Sume lists it in the Video Router catalog, so you call it with the same API key and wallet as every other video model, billed at the provider list price times 1.25.

What stays and what changes

Think of Recast as a person-for-person substitution on a clip you already like. The performance, the framing and the edit are the source; the identity is the variable. That makes it the right tool when a take is good but the casting is wrong, and the wrong tool when you want a different performance, a different location or a different shot list. For those, generate a new clip instead.

Recast behavior as documented, read 2026-10-03
AspectDocumented behaviorSource
People on screenReplaced by the people in your reference photos, one photo per personSume Video Router docs; fal API schema
Motion and cameraKept from the source videofal model page
CutsKept; no single shot in the source may exceed 15 secondsfal API schema; Sume Video Router docs
AudioSource audio is kept in the outputfal API schema
Source length5 to 30 seconds; output duration is the source lengthSume Video Router docs
Photos1 to 4 images, one per personSume Video Router docs
Resolution768p (Sume default) or 1080pSume Video Router docs
PromptOptional, up to 2,000 charactersSume docs/api video-router table

A request that replaces one actor

Send the source and one photo of the replacement. With one person on screen and one photo you do not need a prompt. The call below is the documented Video Router shape; Idempotency-Key makes a retry return the original job instead of a second charge.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: recast-actor-001" \
  -d '{
    "model": "h3-max-recast",
    "video_url": "https://example.com/scene.mp4",
    "reference_image_urls": ["https://example.com/new-actor.jpg"],
    "resolution": "768p",
    "mode": "async"
  }'

What you pay and what can stop the job

Billing is per second of source video. fal lists Recast at $0.30 per output second at 768p and $0.45 at 1080p, with reference images included at no extra charge (read 2026-10-03). Sume's catalog applies a 1.25 multiplier, so a 10-second source costs about $3.75 at 768p and about $5.63 at 1080p before any rounding. For the per-clip arithmetic see what a 12-second clip bills.

Three limits decide whether a clip can be recast at all. The source must be 5 to 30 seconds. No single shot may run longer than 15 seconds, so a locked-off 20-second take is out until you split it. And the model is source-video-only: Sume rejects a text-only request to h3-max-recast, and sume/auto never routes to it, so you must name the id yourself.

  • Public HTTPS URLs only for the video and the photos.
  • One photo per person; photos map to people left to right by default, and the prompt can override that.
  • No aspect_ratio field: output follows the source.

Before you submit a replacement

A short checklist saves a failed job and a retry. First, probe the source: confirm it is between 5 and 30 seconds and look at where the cuts fall, because a shot longer than 15 seconds will not be accepted. The video inspect docs describe a probe plus stills call you can run on a Sume-hosted clip. Second, count the people you want replaced and gather exactly that many photos; the default mapping is left to right across the frame, so order your list the way the cast stands, or say who becomes whom in the prompt. Third, choose a resolution: 768p is the default and the cheaper tier, and a draft at 768p is a sensible way to check the casting before paying for 1080p.

Finally, decide what happens to the voice. The output keeps the source audio, so a replaced actor will still sound like the original performer. That is a feature when the line reading is the thing you love and a problem when the new person should sound different. Sume does not change that in the Recast call; plan the voice as a separate step.

When to pick something else

If the new person should be a face on your own saved avatar and the clip is 4 to 15 seconds with usable audio, the Beta face swap endpoint applies one ready avatar face. If you want a photo to perform a motion from a reference video rather than replace someone in footage, that is Kling motion control. Comparing these is covered in face swap vs Recast. And for a replaced person who must not be a real individual who has not agreed, do not use another person's likeness without permission; Sume's docs do not make a legal determination for you.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume