MiniMax H3 Max lip sync API: a still plus audio, 5 to 14.8 seconds

Sume runs MiniMax H3 Max lip sync at POST /v1/minimax/h3-max/lip-sync: send a still and Sume-hosted audio of 5 to 14.8 seconds. Body, resolutions and price.

4 min readSume
All posts

To make a still image speak with MiniMax H3 Max on Sume, send POST /v1/minimax/h3-max/lip-sync with an image, your audio and its duration_seconds. The audio must be 5 to 14.8 seconds, and the output length follows the audio. It is a separate endpoint from POST /v1/videos, and it does not take a model field.

Facts are from the Models overview docs, the request schema and pricing code, read 2026-09-29.

What does the request take?

The same still-plus-audio body as Sume's other talking-still endpoint. Exactly one visual source is required, and the audio must be hosted by Sume.

curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: h3-lipsync-001" \
  -d '{
    "image_url": "https://example.com/presenter.png",
    "audio_url": "https://media.sume.com/your-workspace-audio.wav",
    "duration_seconds": 9,
    "resolution": "768p"
  }'

What are the rules?

From the Sume request schema and pricing code, read 2026-09-29. Prices are before the 5.5% default agent fee.
FieldRule
image_url or an avatarExactly one visual source
audio_urlSume-hosted, under 10 MB
duration_seconds5 to 14.8; outside that is refused, not clamped
resolution480p, 768p (the default) or 1080p; no 2K
model or endpoint fieldsRefused; Sume selects the model
Price per audio second$0.06 at 480p, $0.10 at 768p, $0.20 at 1080p

Why refuse audio outside 5 to 14.8 seconds?

The provider rejects audio shorter than 5 seconds and silently clips audio past 14.8, which would drop speech. Sume refuses both edges at admission instead. Split longer speech into pieces and join the clips afterward.

When should I use video generation instead?

When you want the model to invent the scene and the sound. With minimax-h3 or minimax-h3-max on POST /v1/videos you can also send an audio reference to steer the voice; see voice reference with audio clips. Lip sync fits when the words are already recorded and must match exactly.

Sources

Related posts

More in Models

All Models posts

Written by Sume