MiniMax H3 Max lip sync API: your first clip in Python
Submit a still and a Sume-hosted audio file to POST /v1/minimax/h3-max/lip-sync, poll the job, and read the result. Python stdlib, with the 5 to 14.8 s rule.

To make a still talk with MiniMax H3 Max on Sume, POST a public HTTPS image (or an avatar handle) and a Sume-hosted audio URL to /v1/minimax/h3-max/lip-sync with a duration_seconds between 5 and 14.8, then poll the returned job until it completes. The API refuses a duration outside that window instead of clamping it, and the audio must be hosted on Sume and no larger than 10 MB.
MiniMax H3 reached the trackers on 2026-07-31 with a 15-second, 2K, native-audio video model and open weights on 2026-08-03 (Magic Hour tracker, read 2026-10-06). Sume's lip-sync route is a separate surface from that video model: per the Sume models docs, read the same day, it is a still-plus-audio endpoint with its own model id, minimax/h3-max/lip-sync.
What does the request look like?
The body is the same as Sume's Fabric route. Send exactly one visual source, image_url or avatar_id or avatar_handle, plus audio_url and duration_seconds. resolution is 480p, 768p (the default) or 1080p; there is no 2K. Do not send model, endpoint or provider_endpoint, because the URL is the model and the body is strict. Add an Idempotency-Key so a retry cannot create a second job.
import json, os, time, urllib.request
BASE = "https://api.sume.com"
KEY = os.environ["SUME_API_KEY"]
def call(method, url, body=None, key=None):
headers = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
if key:
headers["Idempotency-Key"] = key
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(url, data=data, headers=headers, method=method)
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)
def main():
body = {"image_url": os.environ["STILL_URL"], "audio_url": os.environ["AUDIO_URL"],
"duration_seconds": 5.84, "resolution": "768p"}
sub = call("POST", BASE + "/v1/minimax/h3-max/lip-sync", body, "lipsync-first-001")
job_id = sub["job"]["id"]
while True:
st = call("GET", f"{BASE}/v1/jobs/{job_id}/status")
if st.get("status") in ("completed", "failed", "canceled"):
break
time.sleep(5)
print(json.dumps(call("GET", f"{BASE}/v1/jobs/{job_id}/result"), indent=2))
main()What should you set, and what will fail?
Most first-run failures are input rules, not outages.
| Field | Rule | What happens otherwise |
|---|---|---|
| image_url | Public HTTPS still, aspect ratio 0.4 to 2.5 | Rejected before submit |
| audio_url | Sume-hosted, 10 MB or less, 5 to 14.8 s | unsupported_audio_source or audio_too_large |
| duration_seconds | Required, 5 to 14.8, used for the reservation only | invalid_request, never clamped |
| resolution | 480p, 768p (default) or 1080p | 2K is refused |
| Visual source | image_url or avatar_id or avatar_handle, exactly one | Request is rejected |
What does it cost, and when is it done?
Cost is ceil(duration_seconds) times the per-second rate for the resolution. Derived from the docs, that is $0.0625, $0.10 and $0.20 a second at 480p, 768p and 1080p, so the 5-second 768p example reserves $0.50. Sume reserves at admit, captures at completion and refunds on failure. usage.billable_amount_usd_micros in the submit response gives the exact figure for your request.
The job follows the standard lifecycle in Jobs and results: queued, processing, then completed, failed or canceled. For many clips, switch from polling to a webhook so you do not hold a loop open. If your line is longer than 14.8 seconds, cut it at a sentence boundary into parts of at least 5 seconds each; the fresh-audio post shows how to generate audio that lands in the window.
What should you check after the first clip?
Read the job result before you celebrate. A completed job carries a media.sume.com URL for the MP4; open it and check that the mouth is moving to the right syllables at the start and the end, not just the middle. If the audio was shorter than the duration_seconds you sent, you were still reserved on the declared value, but billing follows the ceiling of the real clip duration, so keep the two numbers close.
Next, make the call idempotent. Add an Idempotency-Key that encodes the clip, such as the script id and a revision number, so a network retry after a timeout returns the original job rather than starting a second paid one. If the first attempt fails, the reservation is refunded, and you can resubmit under a new key. Only after that works for one clip should you loop over a list.
Sources
Related posts
More in Media tools
- MiniMax H3 lip sync at 2K? Sume offers 480p, 768p and 1080p
MiniMax H3 the video model lists 2K, but Sume's H3 Max lip-sync route stops at 1080p. The three resolutions, their derived per-second prices and how to choose.
- MiniMax H3 Max video or H3 Max lip sync: fixing the wrong_tool error
generate_video refuses minimax/h3-max/lip-sync with wrong_tool. Use avatar-image-to-video_create for lip sync, minimax-h3-max for text or image video.
- Reddit's 15-second view rule: a lip-sync clip that stays under it
Reddit's Engaged Video Views beta bills videos over 15 s at 15 s. Sume H3 Max lip sync tops out at 14.8 s of audio, so one clip fits. Specs, math and a caveat.
- Reel captions in two languages: language hint vs auto-detect
How Sume video captions pick a language and a style: the optional language hint, auto-detect, the Hangul rule that returns 400, and a Korean Reel example.
Written by Sume