MiniMax H3 Max on Sume: video, lip-sync and recast, which id to call

Three Sume ids carry the MiniMax H3 Max name: minimax-h3-max for video, a lip-sync route and h3-max-recast. What each takes, its window and its rate.

5 min readSume
All posts

There are three different things on Sume with H3 Max in the name, and they take different inputs. minimax-h3-max is a video model on POST /v1/videos that makes a 5 to 15 second clip from text, a frame or references. The H3 Max lip-sync route, POST /v1/minimax/h3-max/lip-sync, makes a talking clip from a still and a 5 to 14.8 second audio file. h3-max-recast replaces the people in a source video with people from 1 to 4 photos. Pick by the asset you already have: a prompt, an audio file or a finished video.

Each of the three has its own request body and its own limits, so read the catalog row for the one you pick and do not carry fields over from another. The most common slip is sending a duration_seconds to lip-sync as if it set the output length. It does not: the audio sets it, and the declared number is used to reserve the price. Recast works the same way, since the source video sets its length.

Plain H3 (minimax-h3) is a fourth id and is covered at the end of the page, so that nobody mixes it up with the first.

The three side by side

All values are from the Sume route and router documentation.

Three H3 Max entry points on Sume (Sume docs, read 2026-10-07)
minimax-h3-maxH3 Max lip-synch3-max-recast
CallPOST /v1/videosPOST /v1/minimax/h3-max/lip-syncVideo Router with video_url
You providePrompt, frames or referencesStill or avatar, plus audio_urlvideo_url plus 1 to 4 reference_image_urls
Length5 to 15 s, you set it5 to 14.8 s, the audio sets it5 to 30 s, the source sets it
Resolution480p, 768p, 1080p480p, 768p, 1080p768p, 1080p
AudioNative stereo, no toggleYour audioSource sound kept
Rate per second$0.0625 / $0.10 / $0.20$0.0625 / $0.10 / $0.20$0.375 / $0.5625

When each one is the right id

Use minimax-h3-max when you are making a new clip. It takes text, a first frame with an optional last frame, or references, and the audio is generated with the clip. Use the lip-sync route when the words are already decided: you bring the voice, and the mouth is made to match. Use Recast when the footage is shot and the person is wrong, for example a presenter who has to change for another market.

A quick test of which one you need is to ask what you can already hold in your hand. If it is only an idea, you need the video model. If it is a recorded voice and a face, you need lip-sync. If it is a finished shot with the wrong person in it, you need Recast. Asking the question in this order avoids the common mistake of trying to fix a finished clip by generating a new one, which gives a different clip and not a corrected one. The three do not substitute for each other. Lip-sync cannot change a shot, Recast cannot make a new scene and the video model cannot keep your audio. A comparison of the talking-photo options is in three routes for a talking photo.

How they bill

The per-second rates are the provider list rate times 1.25. The video model and lip-sync share the 480p, 768p and 1080p tiers, with a list rate of $0.05, $0.08 and $0.16 for video and the same derived prices for lip-sync at the resolutions. Recast has a list rate of $0.30 at 768p and $0.45 at 1080p per second of the source, rounded up. As examples at 768p, a 10 second video clip is $1.00, a 10 second lip-sync is $1.00 and a 10 second Recast job is $3.75.

The full table is in every MiniMax price in one table.

The plain H3 id

minimax-h3 is the standard model with 480p and 768p only, native stereo audio and a 5 to 15 second window. It is the cheaper video option of the family, at a list rate of $0.05 and $0.06 per second. Choosing between it and H3 Max is covered in which id for a 5 to 15 second clip.

Sources

Related posts

More in Models

All Models posts

Written by Sume