Swap the model in a clothing video with AI: H3 Max Recast
Recast swaps the person in a clip for a person from a photo, 5 to 30 seconds. What it keeps, what the docs leave open about the garment, and the Sume call.

To swap the model in a fashion clip, send the clip and one photo per person to h3-max-recast on POST /v1/videos: it replaces each person in the source video with the person in a reference photo, for clips of 5 to 30 seconds. It is a person swap, not a garment swap. The photo says who appears; what that person wears in the output is not something Sume's docs state, so test it on a five-second cut before you commit a lookbook.
The trend behind this is fal's Recast on MiniMax H3 Max, reported on October 1 2026, which Sume ships as h3-max-recast.
What does the model keep?
The Neuron digest, a news aggregator we cite as a report rather than a spec, says fal's Recast swaps people in a source video for people from reference photos while preserving motion, gestures, camera movement, cuts, timing and soundtrack, takes up to four reference images, and handles videos of 5 to 30 seconds. It lists $0.30 per second at 768p and $0.45 at 1080p. Those are fal's reported prices, not Sume's; read Sume's own on the model catalog before you plan a budget.
That list is about the clip, not the clothes. For a fashion piece, the clothes are the product, so what the swapped person wears is the question that matters.
What does Sume's version accept?
From the Video generation docs: h3-max-recast takes one source video plus 1 to 4 person photos, at 768p or 1080p, with an optional prompt, for 5 to 30 seconds. The Video Router page adds that duration is the source length and that each photo is one person. Sume's API code rejects text-only calls, a missing duration, aspect_ratio and generate_audio.
| Field | Value | Note |
|---|---|---|
| model | h3-max-recast | Catalog id |
| input_references | one video_url plus 1 to 4 image_url | One photo per person to swap |
| duration | 5 to 30, whole seconds | Equals the source clip's length |
| resolution | 768p or 1080p | Pick 768p to test |
| prompt | Optional | Say which person becomes which photo |
| aspect_ratio, generate_audio | Not accepted | Output follows the source |
How do you run it?
Host the clip on media.sume.com or at another public HTTPS URL, host the person photo, then submit and poll. The job ends with unsigned_urls[0]. A clip longer than 30 seconds has to be trimmed first; the trim, recast and rejoin post shows how.
Send an Idempotency-Key built from the clip and photo, so a retry does not run twice.
import os
import time
import requests
key = os.environ["SUME_API_KEY"]
h = {"Authorization": f"Bearer {key}"}
ref = lambda k, u: {"type": k, k: {"url": u}}
body = {
"model": "h3-max-recast",
"duration": 10, # whole seconds, 5 to 30, equal to the source clip length
"resolution": "768p",
"prompt": "Person 1 becomes the woman in the photo.",
"input_references": [
ref("video_url", "https://media.sume.com/artifacts/artf_demo/lookbook.mp4"),
ref("image_url", "https://cdn.example.com/new-model.png"),
],
}
job = requests.post("https://api.sume.com/v1/videos", headers={**h, "Idempotency-Key": "recast-look7-v1"}, json=body, timeout=60).json()
while True:
time.sleep(30)
s = requests.get(job["polling_url"], headers=h, timeout=60).json()
if s["status"] == "completed":
print(s["unsigned_urls"][0])
break
if s["status"] in ("failed", "cancelled"):
raise SystemExit(s.get("error", s["status"]))Why not use an image edit instead?
If the thing you want to change is the clothes, an image edit is the more direct tool. A person swap changes who is on screen and leaves the clip's motion, cuts and timing as they were, so it suits one expensive shoot that you want to reuse with several faces. A garment change on a fixed person is a different job, and the outfit change post covers the video-to-video edit for it.
Pick by what changes. A new face in the same shoot is Recast. A new garment on the same face is an edit. A new garment on a new face means two steps, and you check the garment after the second one, not before.
What should you test before a lookbook?
Run one short cut of the clip with the new person's photo and look at the garment frame by frame. If you need a specific outfit on the new person, choose a photo that shows the outfit, then compare the output to your packshot. We cannot tell you from the docs that the outfit transfers, and an unverified claim in a fashion listing is worse than none. Use only photos of people who agreed to appear, and label the clip as AI-edited where a platform asks.
- Test at 768p first; move to 1080p only for the take you will publish.
- Match one reference photo to one person in the source.
- Check hands and fabric edges on moving frames.
- Keep the original clip; the swap does not touch it.
- If you need a different garment on the same person, use an image edit, not Recast.
Sources
Related posts
More in Media tools
- Turn a ChatGPT try-on image into a video with Seedance 2.5
Saved a try-on image to your ChatGPT Library? Host it, then use it as the first frame of a 9:16 clip on seedance-2.5 through POST /v1/videos. A working script.
- Turn a hum into music with AI: what Sume takes as input
Stability says hum-to-steer is coming. Sume's music API takes text and one optional image, not audio. Here is how to describe a hummed tune in a prompt.
- Vertical video subtitles: BBC's 3-line rule and Sume settings
BBC guidance for 9:16 subtitles: up to 3 lines, 90% width, placed a little high. How to set safe_width_ratio and anchor_ratio on Sume burned-in captions.
- Vimeo "Invalid Caption File": build clean WebVTT from Sume
Vimeo rejects a cue that starts at the previous cue's end. Sume sentence segments touch, so trim 1 ms, write UTF-8 WebVTT and upload. Script included.
Written by Sume