Face and body swap video AI: Recast vs Sume's avatar face swap

Sume has two ways to put someone else in a video: H3 Max Recast swaps the person from a photo, Beta Face Swap applies a ready avatar's face. Which to use.

5 min readSume
All posts

If you want the person changed and not only the face, use H3 Max Recast (h3-max-recast on the Video Router). If you want one ready avatar's face on a short clip, use Sume's Beta Avatar Face Swap. The people who search for a face and body swap video AI are usually after the first one, but the second is cheaper to reason about when the same identity must appear in many clips.

The two take different inputs, different clip lengths and different handles, so they are not interchangeable endpoints.

What does each one change?

fal describes Recast as recasting people in videos using reference photos while preserving the source motion, camera work, cuts and audio. Sume's catalog entry says it replaces the main people in the source with the people in 1 to 4 photos, one photo per new person. The unit is the person.

Sume's face swap page says its endpoint applies a ready avatar's face onto a public source video. The unit is the face, and the identity must already exist as an Avatar 1.0 avatar with a handle. Neither page claims anything about how each model treats hair, clothing or body shape, so test both on your own footage before you promise a client either result.

How do the inputs and limits compare?

Both read from Sume's docs, with the fal fields for Recast read 2026-10-03.

  • Recast needs no avatar. Photos are enough, and they can be different on every call.
  • Face swap needs an avatar first, which suits a recurring character or a spokesperson identity.
  • Neither takes aspect_ratio. Face swap has no duration knob; on Recast, duration is only the source length used to quote the hold.
H3 Max Recast vs Beta Avatar Face Swap on Sume, read 2026-10-03
H3 Max RecastAvatar Face Swap (Beta)
EndpointPOST /v1/video-router/generatePOST /v1/models/sume/avatar-face-swap/v1.0/runs
Identity input1 to 4 public photo URLsA ready avatar handle
Source clip5 to 30 s, no shot over 15 sAbout 4 to 15 s with usable audio
PromptOptional, 2000 charactersNot supported
Required choiceResolution 768p or 1080p (optional)quality: standard, plus or max (required)
StatusCatalog modelBeta

How do they price?

Recast is billed per source second: fal's $0.30 at 768p or $0.45 at 1080p, times 1.25 on Sume, so $0.375 or $0.5625 a second. A 10 second clip is $3.75 at 768p.

Face swap prices by quality tier, not per second; the tier numbers are in Sume's face swap cost post. The two cannot be compared per second, so compare per finished clip of the length you actually need.

What is an example of each call?

Recast needs only URLs. A one-person swap on a 10 second clip is model: h3-max-recast, a public video_url, one entry in reference_image_urls, duration: 10 and, if you want it, resolution: 1080p. Nothing else is required, and the prompt can be left out.

Face swap needs the avatar first. You create an Avatar 1.0 identity from a photo, wait until it is ready, then post its avatar_handle, the video_url and a quality to the runs endpoint. Sume's face swap from a photo post walks through that two-step flow. The extra step buys you a reusable identity: the same handle can go onto many clips without re-uploading a photo each time.

What can go wrong with each?

Recast fails fast on shape errors. Five photos, a 31 second clip, a signed URL, generate_audio or aspect_ratio are all rejected before a job is created, and the rejections are listed in the rejected-field post. The provider-side rule that bites later is fal's 15 second shot limit.

Face swap's Beta window is shorter, about 4 to 15 seconds, and it expects usable audio, so a silent clip is a poor fit. Because it is Beta, treat its limits as subject to change and re-read the page before a production run. Both flows are asynchronous jobs: submit, keep the job id, then poll GET /v1/jobs/{id}/status and read /result only when the job is completed.

Which one should I pick?

Pick Recast for a source with one to four people whose whole look should change, when you have photos and no avatar, or when the clip is longer than 15 seconds. Pick face swap when a ready avatar is the identity you are publishing under and the clip fits the 4 to 15 second window. In both cases the person in the photo or avatar has to have agreed. Sume's Terms of Service put likeness permission on the submitter.

For disclosure, treat both outputs as altered video of a real person. TikTok's face swap rules and YouTube's disclosure for face-swapped video are summarised in earlier posts.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume