Voiceover too long for H3 Max lip sync: the 5 to 14.8 s window
MiniMax H3 Max lip sync on Sume takes Sume-hosted audio of 5 to 14.8 seconds, refuses other lengths instead of clamping, and caps the file at 10 MB.

MiniMax H3 Max lip sync on Sume turns one still and one audio file into a talking clip, but the audio must run from 5 to 14.8 seconds and be hosted on Sume, up to 10 MB. Sume refuses duration_seconds outside that window with invalid_request instead of clamping it, because clamping would cut speech without telling you.
What the endpoint takes
Submit to POST /v1/minimax/h3-max/lip-sync, with one visual source: a public HTTPS image_url whose aspect ratio is between 0.4 and 2.5, or a ready avatar_id or avatar_handle. Add an audio_url on media.sume.com and a required duration_seconds. The resolution is 480p, 768p (the default) or 1080p, with no 2K.
This is not the prompt-driven minimax-h3-max of the video catalog. It is the explicit alternative for when the audio fits the provider's window. The docs keep Fabric as the default talking model.
curl -X POST https://api.sume.com/v1/minimax/h3-max/lip-sync \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-lipsync-001" \
-d '{
"avatar_handle": "yura",
"audio_url": "https://media.sume.com/artifacts/artf_demo/line.wav",
"duration_seconds": 9.2,
"resolution": "768p"
}'Why the window is strict
The provider rejects audio under 5 seconds and silently clips audio after 14.8 seconds. If Sume clamped your reservation, you would get a clip shorter than your line, and part of the speech would be gone. So the API checks duration_seconds itself and refuses anything outside the window.
| Input | Rule | Error |
|---|---|---|
| duration_seconds | 5 to 14.8, never clamped | invalid_request |
| audio_url | Sume-hosted, 10 MB at most | unsupported_audio_source, audio_too_large |
| Visual source | image_url or avatar, not both | Same as Fabric |
| Body | Strict: no model or endpoint fields | Rejected |
What to do with a longer script
Cut the script into lines of 14 seconds or less, generate each voice line, and submit one lip-sync clip per line. Join the clips later on a Timeline render. You have to do the cutting yourself, because the API does not trim audio for you.
If you need a single take longer than 15 seconds, use the avatar talking-video route instead, which plans 4 to 60 seconds.
Sources
Related posts
More in Sume Avatar 1.0
- What Sume Avatar 1.0 Does and Does Not Do vs Live Avatars
A plain list of what Sume Avatar 1.0 renders (scripted 4-60 s clips) and what it does not do (real-time conversation), set beside Tavus Griffin's live model.
- Can I use Tavus Griffin-Lite yet? Preview status and what to ship
Tavus Griffin-Lite is a research preview for select trusted testers, not open to customers. Here is what the page says, and what you can build on Sume today.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume