MiniMax H3 Max Recast API: swap people in a video, fal price vs Sume
H3 Max Recast swaps people in a source video for reference photos, keeping motion, cuts and audio. fal lists $0.30 a second at 768p; what Sume accepts.

H3 Max Recast replaces the people in an existing video with people from reference photos while keeping the source motion, camera, cuts and audio. fal lists it at $0.30 per second at 768p and $0.45 per second at 1080p, with reference images at no extra charge. Sume exposes the same capability as h3-max-recast: one source video_url, one to four reference_image_urls, 768p or 1080p, and a source of 5 to 30 seconds.
fal's details come from its H3 Max Recast page, read 2026-10-01. Sume's come from the Video Generation and Video Router docs.
What does Recast take in and give back?
fal's description is that Recast changes the people in a clip using reference photos and preserves what the clip does around them. The input is a video URL (mp4, mov, webm, m4v or gif) and multiple reference image URLs, and the output resolution is 768p or 1080p. Sume's docs add the shape of the request: one photo per person to be swapped in, an optional prompt, and an output length that equals the source length.
| Item | fal | Sume `h3-max-recast` |
|---|---|---|
| Source video | Required; mp4, mov, webm, m4v, gif | Required video_url; 5 to 30 seconds, no shot over 15 seconds |
| Reference photos | Multiple, no extra charge | 1 to 4 reference_image_urls, one per person |
| Resolution | 768p or 1080p | 768p (default) or 1080p |
| Prompt | Optional settings | Optional |
| Price per second | $0.30 at 768p, $0.45 at 1080p | List x 1.25: $0.375 at 768p, $0.5625 at 1080p |
What does a 20-second recast cost?
Billing follows the source length: Sume's catalog says duration must be the inspected source length in seconds, rounded up. A 20-second source at 768p is $6.00 on fal and, applying Sume's list x 1.25 rate, $7.50 on Sume. At 1080p it is $9.00 on fal and $11.25 on Sume. Check pricing_skus for the live rate before you run a long file, because Sume reserves the full amount at submit.
When is Recast the wrong tool?
This id needs a source video and reference photos, so it is not a text-only generator. If you want to change an object, a background or a style rather than a person, use the video_url edit mode on gemini-omni-flash-1.1 instead. If your source has a shot longer than 15 seconds, split it first; Sume's catalog lists no single shot longer than 15 seconds as a constraint on the source.
- Use clear, front-facing reference photos, one per person.
- Keep the people count in the clip equal to the photo count.
- Check that you have the right to use each person's likeness before you publish.
Sources
Related posts
More in Models
- MiniMax H3 sound design prompts: direct the audio like the picture
fal's H3 guide says to direct audio as deliberately as picture: name sonic elements, not 'music'. What it looks like in a Sume request, and what you can't set.
- Where are the lyrics in a Lyria 3.5 result? Gemini vs Sume
Google returns Lyria 3.5 lyrics and song structure as text beside the audio. Sume puts model-reported lyrics or a section map in result.lyrics when present.
- Nano Banana negative prompt: describe what you want instead
Google's Gemini image guide says to write semantic negative prompts: describe an empty street, not 'no cars'. Sume's image request has no negative field.
- Nano Banana prompt languages: Korean and Japanese, per Google
Google lists the languages Gemini image models work best in, including ko-KR and ja-JP. What that means for a Korean or Japanese prompt sent through Sume.
Written by Sume