Face swap vs H3 Max Recast: which swaps the person in a video?

Sume offers two ways to put a different person in a video: Avatar Face Swap (Beta) and H3 Max Recast. Inputs, length limits, audio and price side by side.

5 min readSume
All posts

Use Face Swap (Beta) when you want the same recurring character in a short clip, up to about 15 seconds. Use H3 Max Recast when you have one to four reference photos, a clip up to 30 seconds, or more than one person to swap. Face Swap works from a saved avatar handle; Recast works from photos you send each time.

Recast landed in the news this week as fal's new person-swap model on MiniMax H3 Max, and Sume lists it as the video model h3-max-recast. Both run on Sume, so the choice is about inputs and length. Facts below come from Sume's docs, its API reference and fal's own model page, read 2026-10-03.

How do the two compare on inputs and limits?

Avatar Face Swap (Beta) vs H3 Max Recast on Sume, read 2026-10-03
Face Swap (Beta)H3 Max Recast
Identity inputA ready avatar handle (create once, $0.95)1-4 reference photos, one per person
People swappedThe avatar replaces the personUp to four people, one photo each
Source lengthAbout 4-15 seconds5-30 seconds
PromptNot acceptedOptional
Quality or resolutionquality required: standard, plus or maxResolution chosen: 768p or 1080p
RoutePOST /v1/models/sume/avatar-face-swap/v1.0/runsPOST /v1/video-router/generate, model h3-max-recast

What happens to the audio?

Face Swap keeps the source audio: Sume's workflow re-attaches the original track to the finished clip, and the Beta expects usable speech. fal says Recast preserves audio, motion, camera work and cuts. On Sume, Recast rejects generate_audio and audio references, so do not plan on adding a new voice there either. Run video inspect on a result to confirm it has sound.

What does each cost?

Face Swap reserves a 15-second maximum at the Avatar Video rate for your tier: $2.76 on standard, $3.68 on plus, $8.25 on max. fal lists Recast at $0.30 per second of output at 768p and $0.45 at 1080p, with reference images included; fal's page is a vendor list price. Sume's Video Router docs say every model bills at list × 1.25, which would put 768p near $0.375 per second by that rule; check the pricing page for the live rate. A 30-second Recast clip is also outside Face Swap's range entirely.

Which should I pick?

  • One brand character, many short clips: Face Swap, because the handle is reusable and consistent.
  • A clip longer than 15 seconds: Recast, up to 30 seconds.
  • Two or three people to replace: Recast, one photo each.
  • You want a prompt to steer the result: Recast, where a prompt is optional.
  • You want the swap to follow your avatar's look including outfit and hair: Face Swap, which targets the whole avatar.
  • You do not have a saved avatar and want a one-off: Recast, which skips the avatar step.

How do the requests differ?

Face Swap is one purpose-built endpoint with a strict body: avatar_handle, video_url and quality, plus the standard mode, webhook_url and wait_timeout_seconds communication fields. Anything else, such as a prompt, an avatar id or a duration, is rejected. Recast goes through the Video Router with model set to h3-max-recast, a video_url and one reference_image_urls entry per person, and its duration is the source's own length.

Both are asynchronous. Each returns job URLs on submit, and you poll the status URL, then fetch the result when result_ready is true. A sync wait is capped at 30 seconds, and a swap will almost always outlast it, so submit async or with a webhook rather than blocking on the response.

Which should I pick for which job?

  • You have a saved avatar and a short clip with speech you want to keep: Face Swap. The avatar handle is the identity, and the source audio comes back.
  • You have one to four reference photos and want a different person in a 5 to 30 second clip: Recast. The identity comes from the photos, not from a saved handle.
  • You need the original words exactly: Face Swap carries the source audio. Sume rejects generate_audio and audio references on Recast, so plan to handle sound yourself.
  • You need to reuse the same person across many clips: the avatar handle in Face Swap is built for that.

Can I try both on the same clip?

Yes, and for a one-off it is cheap to compare. Face Swap needs a public HTTPS clip of about 4 to 15 seconds and a ready avatar. Recast goes through the Video Router with a media input and your reference photos. Run each once, then compare face stability, mouth sync and audio. Review the results side by side before you spend on a longer clip, and keep the idempotency keys different so neither request is mistaken for a retry of the other.

What is the same for both?

Both need a source you have the right to re-cast, and both produce a synthetic person, so disclose them as AI-generated where the platform requires it. Neither is a live feature; each is a queued job you submit and poll. For Genjutsu, the other person-swap row, see Recast vs Genjutsu.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume