Put a person in a new scene: H3 Max reference video or Recast?
Recast swaps people in a video you have. Reference-to-video makes a new clip from photos. Inputs and 10-second prices: $1.00 vs $3.75 at 768p.

If a video already exists and you only want a different person in it, use H3 Max Recast; if you want a brand-new scene that features a person from a photo, use H3 Max reference-to-video. Recast keeps the source motion, camera, cuts and sound, while reference-to-video invents all of those from a prompt.
On Sume, a 10-second Recast at 768p bills $3.75 and a 10-second minimax-h3-max reference-to-video clip at 768p bills $1.00, but only Recast starts from your footage.
What each route takes
The h3-max-recast row takes one video_url and 1 to 4 reference_image_urls, one photo for each new person. The prompt is optional (up to 2000 characters). There is no text-only generation, and the source is 5 to 30 seconds with no shot longer than 15 seconds. Resolution is 768p by default or 1080p.
The minimax-h3-max row takes a prompt and, for reference-to-video, reference_*_urls. The docs say it accepts image, video and audio references, runs at 480p, 768p or 1080p for 5 to 15 seconds, and always produces native stereo audio. The new clip has whatever camera, cuts and motion the model produces.
Price comparison
Recast bills the source length at $0.375 per second at 768p or $0.5625 at 1080p. H3 Max bills $0.10 per second at 768p or $0.20 at 1080p. Max reference-to-video is not billed per reference image on Sume.
| Length | Recast 768p | Recast 1080p | H3 Max r2v 768p | H3 Max r2v 1080p |
|---|---|---|---|---|
| 5 s | $1.88 | $2.82 | $0.50 | $1.00 |
| 10 s | $3.75 | $5.63 | $1.00 | $2.00 |
| 15 s | $5.63 | $8.44 | $1.50 | $3.00 |
What each can and cannot keep
Recast is for continuity: the scene, the choreography, the lighting and the soundtrack come from your source. It is the right tool for swapping a model, a spokesperson or a creator across market versions of the same ad. It cannot create motion you did not film, and a shot over 15 seconds must be split first.
Reference-to-video is for invention: you describe a scene and give the model photos of who or what should appear. It is not an editor, so the output will not match an existing shot. Photos guide appearance but do not lock it, which is why you may need several takes.
- Have a finished video and a new face: Recast.
- Have photos and a script idea, no footage: reference-to-video.
- Need the same background and camera across versions: Recast only.
Cost of getting it right
Because r2v is cheaper per second, trying five takes at 10 seconds is $5.00 at 768p, which is still above one Recast take at $3.75 but below two. If you can film or license a source, Recast tends to give a more predictable result per dollar. If you cannot, reference-to-video is the only route of the two that works without footage. The submit response of each carries usage.billable_amount_usd_micros, so log it to compare takes.
Checking the catalog before you build
The two rows live in the same catalog, but their capabilities differ, and the docs tell callers to read capabilities from GET /v1/video-router/models and not to assume one envelope. For Recast, look at the duration range (5 to 30 seconds), the resolutions (768p, 1080p) and the reference photo count (1 to 4). For H3 Max, look at supported_input_references and the resolution list.
Neither route is selected by sume/auto. The Video Router docs say callers select h3-max-recast explicitly, and that Sume's own routing never picks it, so a request that names the id is the only way to reach it.
A practical rule of thumb
Ask what you already have. If you have a finished video whose motion, framing and timing you like, and you only want a different person in it, use Recast: the source video supplies the performance and you supply one to four photos. If you have only photos and a prompt, use reference-to-video on minimax-h3-max, where the model invents the scene and the motion.
The cost shapes the decision too. A 10-second H3 Max take at 768p is $1.00 and can be rerun cheaply. A Recast of a 10-second source at 768p is billed on the source seconds, so a bad source costs you the same as a good one. Check the source for shots over 15 seconds before you submit, since the docs reject them.
Sources
Related posts
More in Comparisons
- Luma Ray 3.2 features and the Sume endpoint for each one
Sume does not run Ray 3.2. Match its keyframes, reframe, face tracking and 20 s clips to the Sume surfaces that exist, with docs-verified limits.
- Ray 3.2 takes 16 keyframes; Sume frame_images takes two
Luma lists up to 16 keyframes per Ray 3.2 clip. Sume's frame_images array uses first_frame and last_frame. Which catalog models list them, per the docs.
- Ray 3.2 tracks 8 faces; Sume swaps 1-4 people with H3 Max Recast
Luma Ray 3.2 lists facial tracking for up to 8 faces. Sume has no tracking output, but h3-max-recast swaps 1-4 people in a clip and Kling drives a still.
- Reference limits per request, vendor vs Sume: Wan 3.0, Seedance, Omni
What each vendor says one video request can take, next to what Sume's catalog allows: Wan 3.0, Seedance 2.5, Gemini Omni 1.1 Flash, and Ray3.2.
Written by Sume