Face and body swap video AI: Recast vs Sume's avatar face swap
Sume has two ways to put someone else in a video: H3 Max Recast swaps the person from a photo, Beta Face Swap applies a ready avatar's face. Which to use.
If you want the person changed and not only the face, use H3 Max Recast (h3-max-recast on the Video Router). If you want one ready avatar's face on a short clip, use Sume's Beta Avatar Face Swap. The people who search for a face and body swap video AI are usually after the first one, but the second is cheaper to reason about when the same identity must appear in many clips.
The two take different inputs, different clip lengths and different handles, so they are not interchangeable endpoints.
What does each one change?
fal describes Recast as recasting people in videos using reference photos while preserving the source motion, camera work, cuts and audio. Sume's catalog entry says it replaces the main people in the source with the people in 1 to 4 photos, one photo per new person. The unit is the person.
Sume's face swap page says its endpoint applies a ready avatar's face onto a public source video. The unit is the face, and the identity must already exist as an Avatar 1.0 avatar with a handle. Neither page claims anything about how each model treats hair, clothing or body shape, so test both on your own footage before you promise a client either result.
How do the inputs and limits compare?
Both read from Sume's docs, with the fal fields for Recast read 2026-10-03.
- Recast needs no avatar. Photos are enough, and they can be different on every call.
- Face swap needs an avatar first, which suits a recurring character or a spokesperson identity.
- Neither takes
aspect_ratio. Face swap has no duration knob; on Recast,durationis only the source length used to quote the hold.
| H3 Max Recast | Avatar Face Swap (Beta) | |
|---|---|---|
| Endpoint | POST /v1/video-router/generate | POST /v1/models/sume/avatar-face-swap/v1.0/runs |
| Identity input | 1 to 4 public photo URLs | A ready avatar handle |
| Source clip | 5 to 30 s, no shot over 15 s | About 4 to 15 s with usable audio |
| Prompt | Optional, 2000 characters | Not supported |
| Required choice | Resolution 768p or 1080p (optional) | quality: standard, plus or max (required) |
| Status | Catalog model | Beta |
How do they price?
Recast is billed per source second: fal's $0.30 at 768p or $0.45 at 1080p, times 1.25 on Sume, so $0.375 or $0.5625 a second. A 10 second clip is $3.75 at 768p.
Face swap prices by quality tier, not per second; the tier numbers are in Sume's face swap cost post. The two cannot be compared per second, so compare per finished clip of the length you actually need.
What is an example of each call?
Recast needs only URLs. A one-person swap on a 10 second clip is model: h3-max-recast, a public video_url, one entry in reference_image_urls, duration: 10 and, if you want it, resolution: 1080p. Nothing else is required, and the prompt can be left out.
Face swap needs the avatar first. You create an Avatar 1.0 identity from a photo, wait until it is ready, then post its avatar_handle, the video_url and a quality to the runs endpoint. Sume's face swap from a photo post walks through that two-step flow. The extra step buys you a reusable identity: the same handle can go onto many clips without re-uploading a photo each time.
What can go wrong with each?
Recast fails fast on shape errors. Five photos, a 31 second clip, a signed URL, generate_audio or aspect_ratio are all rejected before a job is created, and the rejections are listed in the rejected-field post. The provider-side rule that bites later is fal's 15 second shot limit.
Face swap's Beta window is shorter, about 4 to 15 seconds, and it expects usable audio, so a silent clip is a poor fit. Because it is Beta, treat its limits as subject to change and re-read the page before a production run. Both flows are asynchronous jobs: submit, keep the job id, then poll GET /v1/jobs/{id}/status and read /result only when the job is completed.
Which one should I pick?
Pick Recast for a source with one to four people whose whole look should change, when you have photos and no avatar, or when the clip is longer than 15 seconds. Pick face swap when a ready avatar is the identity you are publishing under and the clip fits the 4 to 15 second window. In both cases the person in the photo or avatar has to have agreed. Sume's Terms of Service put likeness permission on the submitter.
For disclosure, treat both outputs as altered video of a real person. TikTok's face swap rules and YouTube's disclosure for face-swapped video are summarised in earlier posts.
Sources
Related posts
More in Comparisons
- Face swap vs H3 Max Recast: which swaps the person in a video?
Sume offers two ways to put a different person in a video: Avatar Face Swap (Beta) and H3 Max Recast. Inputs, length limits, audio and price side by side.
- FLUX 3 Image alternatives on Sume, feature by feature
No FLUX 3 Image in Sume's catalog yet. Match each FLUX 3 feature, 10 references, 4K, region edits, grounding, to the Sume model that has it or the gap.
- FLUX 3 Image vs Nano Banana Pro for 4K editing on Sume
FLUX 3 Image is not in Sume's catalog; Nano Banana Pro is, with a 4K tier and 10 reference images. A checklist of what each does for editing, from vendor pages.
- Full-duplex video AI: what Griffin changes, what Sume does
Full-duplex video AI listens, watches and answers at once. Tavus Griffin is gated; Sume makes scripted avatar clips by job. Where each fits.
Written by Sume