MiniMax H3 lip sync at 2K? Sume offers 480p, 768p and 1080p

MiniMax H3 the video model lists 2K, but Sume's H3 Max lip-sync route stops at 1080p. The three resolutions, their derived per-second prices and how to choose.

5 min readSume
All posts

Sume's MiniMax H3 Max lip-sync route does not produce 2K. Its resolution field accepts 480p, 768p (the default) and 1080p, and a 2K request is rejected. The 2K you may have read about belongs to the MiniMax H3 video model, which a tracker lists as a 15-second, 2K, native stereo audio release dated 2026-07-31 (Magic Hour tracker, read 2026-10-06).

That is easy to mix up because Sume lists both. The video model appears in the video catalog as minimax-h3 and minimax-h3-max for text, frame and reference prompts, and the lip-sync route has its own id, minimax/h3-max/lip-sync, for still-plus-audio talking clips (Video generation docs, Models, both read 2026-10-06). They are priced and validated separately.

What does each resolution cost?

Sume bills ceil(duration_seconds) times the per-second list rate for the resolution times a 1.25 margin. The list rates in the docs are $0.05, $0.08 and $0.16 a second.

Derived from the MiniMax H3 Max Lip Sync section of the Sume models docs (list rate x 1.25), read 2026-10-06.
ResolutionPer second5 s clip10 s clip14.8 s audio, billed as 15 s
480p$0.0625$0.3125$0.625$0.9375
768p (default)$0.10$0.50$1.00$1.50
1080p$0.20$1.00$2.00$3.00

How do you choose?

Resolution is separate from length. The audio sets the output duration, and the window is 5 to 14.8 seconds, so a 14.8-second line is billed as 15. A shorter line costs less, but audio under 5 seconds is outside the provider's window, and Sume refuses a duration_seconds outside 5 to 14.8 rather than clamping it. Route those short lines to Fabric instead of padding or re-cutting the audio.

  • Pick 480p for drafts and lip-accuracy checks. It is the cheapest way to hear whether the mouth tracks your audio.
  • Pick 768p for most social and web placements. It is the default and the figure the doc's 5-second reservation example uses.
  • Pick 1080p only when the clip will be viewed full-frame on a large screen, and compare it side by side first; it costs twice the 768p rate.
  • Never plan on 2K here. If you need a 2K talking clip, this route is not the way to get it.

How should the clip join the rest of an edit?

Use one lip-sync model for a whole run. Sume's guidance is that the model sets the frame rate of the assembled output and that you should measure H3 Max on your first real clip; the frame-rate post shows how. Mixing two lip-sync models in one timeline invites a frame-rate mismatch you only see on playback.

For a complete price list across the MiniMax ids on Sume, including the video model and Recast, see every MiniMax price in one table. If you are choosing between this route and Sume's default talk model, Fabric, that is a quality-and-price decision to settle on a sample: run the same 10-second audio through both at their default resolution and keep whichever you cannot tell from a recording of the same person.

How do you pick a tier without guessing?

Render the same 5-second line at all three resolutions once, which costs $0.3125, $0.50 and $1.00 at the documented Sume rates, and look at each on the device where it will be watched. A phone feed often hides the gap between 480p and 768p; a landing-page hero usually does not. Write down the winner per placement and make it a constant in your code so the choice is reviewed, not drifted.

Remember that the 2K output in the MiniMax H3 release notes belongs to the video model, not to this lip-sync route. Sending a 2K value here is an invalid request. If a deliverable needs more than 1080p, lip sync is not the route to use for that asset, and the right move is to say so early rather than upscale and hope.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume