Talking photo: H3 Max lip sync, Kling motion control or Recast?

Lip sync follows your audio, motion control follows a reference video, Recast swaps a person. 10-second prices on Sume: $1.00, $1.58 and $3.75 at 768p.

5 min readSume
All posts

If you have a still and an audio line, use MiniMax H3 Max lip sync: it takes a 5 to 14.8 second audio clip and a photo and bills per second of audio. If you have a still and a performance video, use Kling 3.0 motion control. If you have a finished video and want a different person in it, use H3 Max Recast.

The three routes look alike in a catalog because each one starts from a still or a short source, but their inputs, price per second and lengths are different. This post puts them in one table.

What each route takes

H3 Max lip sync (POST /v1/minimax/h3-max/lip-sync) takes image_url or a ready avatar, plus a Sume-hosted audio_url of up to 10 MB and 5 to 14.8 seconds. duration_seconds is required and the API rejects values outside 5 to 14.8 instead of clamping them. Resolution is 480p, 768p (default) or 1080p, with no 2K.

Kling 3.0 motion control (POST /v1/kling/3.0/motion-control) takes image_url or an avatar and a motion_video_url that drives all the output motion. duration_seconds is 1 to 30 and is used for the reservation only, because the motion video sets the real length. keep_original_sound defaults to true.

H3 Max Recast is a Video Router row, h3-max-recast: one source video_url and 1 to 4 reference photos, 5 to 30 seconds, no shot over 15 seconds. The output keeps the source motion, camera, cuts and sound.

Price for a 10-second result

Lip sync bills ceil(duration_seconds) times the fal list per second times 1.25, with list prices of $0.05, $0.08 and $0.16 for 480p, 768p and 1080p. Kling motion control bills $0.126 per output second times 1.25, which is $0.1575. Recast bills $0.30 or $0.45 per second at list, times 1.25.

10-second result, Sume billable at list x 1.25 (repo docs read 2026-10-05)
RouteInput you supplyPer second10 s
H3 Max lip sync 768pPhoto + audio$0.10$1.00
H3 Max lip sync 1080pPhoto + audio$0.20$2.00
Kling motion controlPhoto + motion video$0.1575$1.58 (rounded up)
H3 Max Recast 768pSource video + person photos$0.375$3.75
H3 Max Recast 1080pSource video + person photos$0.5625$5.63 (rounded up)

Which one answers which question

Pick lip sync when the words are already written and recorded, for example a text-to-speech line. Sume's own docs say talking shots go through text-to-speech first and then a lip-sync model, and that H3 Max lip sync is the explicit alternative when the audio fits the 5 to 14.8 second window, with Fabric as the default talk model.

Pick motion control when the thing you have is a movement, such as a dance or a gesture reference, and you want your character to perform it. The sound comes from the reference video when keep_original_sound is true, so it does not make a photo speak your script.

Pick Recast when the video exists and the person is the thing to change. Do not use it to animate a still; it has no text-only or still-only generation.

  • Audio under 5 s does not fit H3 Max lip sync; use Fabric.
  • A reference longer than 30 s does not fit motion control; trim it first.
  • A shot longer than 15 s does not fit Recast; split it first.

A note on cost control

All three routes reserve the amount at admit, capture it on completion and refund it on failure, per the Sume docs. Read usage.billable_amount_usd_micros from each submit response so you see the quote before the render finishes.

Three quick scenarios

A support team wants a 12-second spoken greeting from a brand photo. The line is already text, so they generate speech first and send the audio to H3 Max lip sync at 768p: 12 seconds x $0.10 = $1.20. If the line were 4 seconds long it would not fit the 5 to 14.8 second window, and Fabric is the route the docs name for those segments.

A creator wants a cartoon fox to repeat a dance from a phone video. That is motion control: a 20-second reference at $0.1575 per second is $3.15, and the fox is the character from the still. The creator does not need to supply audio, since the reference sound is kept by default.

A brand has a 25-second ad with an actor and wants a different creator in it. That is Recast: 25 seconds x $0.375 = $9.38 at 768p after rounding up, with the original camera, cuts and sound preserved.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume