Voiceover too long for H3 Max lip sync: the 5 to 14.8 s window

MiniMax H3 Max lip sync on Sume takes Sume-hosted audio of 5 to 14.8 seconds, refuses other lengths instead of clamping, and caps the file at 10 MB.

5 min readSume
All posts

MiniMax H3 Max lip sync on Sume turns one still and one audio file into a talking clip, but the audio must run from 5 to 14.8 seconds and be hosted on Sume, up to 10 MB. Sume refuses duration_seconds outside that window with invalid_request instead of clamping it, because clamping would cut speech without telling you.

What the endpoint takes

Submit to POST /v1/minimax/h3-max/lip-sync, with one visual source: a public HTTPS image_url whose aspect ratio is between 0.4 and 2.5, or a ready avatar_id or avatar_handle. Add an audio_url on media.sume.com and a required duration_seconds. The resolution is 480p, 768p (the default) or 1080p, with no 2K.

This is not the prompt-driven minimax-h3-max of the video catalog. It is the explicit alternative for when the audio fits the provider's window. The docs keep Fabric as the default talking model.

curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: h3-lipsync-001" \
  -d '{
    "avatar_handle": "yura",
    "audio_url": "https://media.sume.com/artifacts/artf_demo/line.wav",
    "duration_seconds": 9.2,
    "resolution": "768p"
  }'

Why the window is strict

The provider rejects audio under 5 seconds and silently clips audio after 14.8 seconds. If Sume clamped your reservation, you would get a clip shorter than your line, and part of the speech would be gone. So the API checks duration_seconds itself and refuses anything outside the window.

From the Sume H3 Max lip sync docs, read 2026-10-05
InputRuleError
duration_seconds5 to 14.8, never clampedinvalid_request
audio_urlSume-hosted, 10 MB at mostunsupported_audio_source, audio_too_large
Visual sourceimage_url or avatar, not bothSame as Fabric
BodyStrict: no model or endpoint fieldsRejected

What to do with a longer script

Cut the script into lines of 14 seconds or less, generate each voice line, and submit one lip-sync clip per line. Join the clips later on a Timeline render. You have to do the cutting yourself, because the API does not trim audio for you.

If you need a single take longer than 15 seconds, use the avatar talking-video route instead, which plans 4 to 60 seconds.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume