MiniMax H3 Max lip sync API: a still plus audio, 5 to 14.8 seconds
Sume runs MiniMax H3 Max lip sync at POST /v1/minimax/h3-max/lip-sync: send a still and Sume-hosted audio of 5 to 14.8 seconds. Body, resolutions and price.

To make a still image speak with MiniMax H3 Max on Sume, send POST /v1/minimax/h3-max/lip-sync with an image, your audio and its duration_seconds. The audio must be 5 to 14.8 seconds, and the output length follows the audio. It is a separate endpoint from POST /v1/videos, and it does not take a model field.
Facts are from the Models overview docs, the request schema and pricing code, read 2026-09-29.
What does the request take?
The same still-plus-audio body as Sume's other talking-still endpoint. Exactly one visual source is required, and the audio must be hosted by Sume.
curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-lipsync-001" \
-d '{
"image_url": "https://example.com/presenter.png",
"audio_url": "https://media.sume.com/your-workspace-audio.wav",
"duration_seconds": 9,
"resolution": "768p"
}'What are the rules?
| Field | Rule |
|---|---|
image_url or an avatar | Exactly one visual source |
audio_url | Sume-hosted, under 10 MB |
duration_seconds | 5 to 14.8; outside that is refused, not clamped |
resolution | 480p, 768p (the default) or 1080p; no 2K |
model or endpoint fields | Refused; Sume selects the model |
| Price per audio second | $0.06 at 480p, $0.10 at 768p, $0.20 at 1080p |
Why refuse audio outside 5 to 14.8 seconds?
The provider rejects audio shorter than 5 seconds and silently clips audio past 14.8, which would drop speech. Sume refuses both edges at admission instead. Split longer speech into pieces and join the clips afterward.
When should I use video generation instead?
When you want the model to invent the scene and the sound. With minimax-h3 or minimax-h3-max on POST /v1/videos you can also send an audio reference to steer the voice; see voice reference with audio clips. Lip sync fits when the words are already recorded and must match exactly.
Sources
Related posts
More in Models
- MiniMax H3 open weights on Hugging Face: what is in the release
MiniMax H3 has a model card on Hugging Face: two checkpoints, a 33B-parameter Transformer, a community license and a 4-GPU serving example. What to check first.
- AI video editing with a text prompt: MiniMax H3 and Sume's edit path
MiniMax H3 is described as editing existing video from instructions. What the vendor says, what Sume's H3 ids accept, and the id with an edit field.
- Nano Banana 2 Lite API: what Google shipped and what Sume lists
Nano Banana 2 Lite is Google's gemini-3.1-flash-lite-image model. What Google says it does, and which Nano Banana models Sume's image API lists today.
- Nano Banana Pro vs Nano Banana 2: which id to send on Sume
Nano Banana Pro and Nano Banana 2 share tiers and 10 references on Sume; they differ in extra aspect ratios and in which tiers change the price.
Written by Sume