MiniMax H3 Max lip sync API: your first clip in Python

Submit a still and a Sume-hosted audio file to POST /v1/minimax/h3-max/lip-sync, poll the job, and read the result. Python stdlib, with the 5 to 14.8 s rule.

5 min readSume
All posts

To make a still talk with MiniMax H3 Max on Sume, POST a public HTTPS image (or an avatar handle) and a Sume-hosted audio URL to /v1/minimax/h3-max/lip-sync with a duration_seconds between 5 and 14.8, then poll the returned job until it completes. The API refuses a duration outside that window instead of clamping it, and the audio must be hosted on Sume and no larger than 10 MB.

MiniMax H3 reached the trackers on 2026-07-31 with a 15-second, 2K, native-audio video model and open weights on 2026-08-03 (Magic Hour tracker, read 2026-10-06). Sume's lip-sync route is a separate surface from that video model: per the Sume models docs, read the same day, it is a still-plus-audio endpoint with its own model id, minimax/h3-max/lip-sync.

What does the request look like?

The body is the same as Sume's Fabric route. Send exactly one visual source, image_url or avatar_id or avatar_handle, plus audio_url and duration_seconds. resolution is 480p, 768p (the default) or 1080p; there is no 2K. Do not send model, endpoint or provider_endpoint, because the URL is the model and the body is strict. Add an Idempotency-Key so a retry cannot create a second job.

import json, os, time, urllib.request

BASE = "https://api.sume.com"
KEY = os.environ["SUME_API_KEY"]

def call(method, url, body=None, key=None):
    headers = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
    if key:
        headers["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body is not None else None
    req = urllib.request.Request(url, data=data, headers=headers, method=method)
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)

def main():
    body = {"image_url": os.environ["STILL_URL"], "audio_url": os.environ["AUDIO_URL"],
            "duration_seconds": 5.84, "resolution": "768p"}
    sub = call("POST", BASE + "/v1/minimax/h3-max/lip-sync", body, "lipsync-first-001")
    job_id = sub["job"]["id"]
    while True:
        st = call("GET", f"{BASE}/v1/jobs/{job_id}/status")
        if st.get("status") in ("completed", "failed", "canceled"):
            break
        time.sleep(5)
    print(json.dumps(call("GET", f"{BASE}/v1/jobs/{job_id}/result"), indent=2))

main()

What should you set, and what will fail?

Most first-run failures are input rules, not outages.

Rules from the MiniMax H3 Max Lip Sync section of the Sume models docs, read 2026-10-06.
FieldRuleWhat happens otherwise
image_urlPublic HTTPS still, aspect ratio 0.4 to 2.5Rejected before submit
audio_urlSume-hosted, 10 MB or less, 5 to 14.8 sunsupported_audio_source or audio_too_large
duration_secondsRequired, 5 to 14.8, used for the reservation onlyinvalid_request, never clamped
resolution480p, 768p (default) or 1080p2K is refused
Visual sourceimage_url or avatar_id or avatar_handle, exactly oneRequest is rejected

What does it cost, and when is it done?

Cost is ceil(duration_seconds) times the per-second rate for the resolution. Derived from the docs, that is $0.0625, $0.10 and $0.20 a second at 480p, 768p and 1080p, so the 5-second 768p example reserves $0.50. Sume reserves at admit, captures at completion and refunds on failure. usage.billable_amount_usd_micros in the submit response gives the exact figure for your request.

The job follows the standard lifecycle in Jobs and results: queued, processing, then completed, failed or canceled. For many clips, switch from polling to a webhook so you do not hold a loop open. If your line is longer than 14.8 seconds, cut it at a sentence boundary into parts of at least 5 seconds each; the fresh-audio post shows how to generate audio that lands in the window.

What should you check after the first clip?

Read the job result before you celebrate. A completed job carries a media.sume.com URL for the MP4; open it and check that the mouth is moving to the right syllables at the start and the end, not just the middle. If the audio was shorter than the duration_seconds you sent, you were still reserved on the declared value, but billing follows the ceiling of the real clip duration, so keep the two numbers close.

Next, make the call idempotent. Add an Idempotency-Key that encodes the clip, such as the script id and a revision number, so a network retry after a timeout returns the original job rather than starting a second paid one. If the first attempt fails, the reservation is refunded, and you can resubmit under a new key. Only after that works for one clip should you loop over a list.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume