How long does AI lip sync take? fal says about a minute at 1080p

fal says a 1080p H3 Max lip-sync clip takes about a minute. On Sume the call is an async job, so poll with backoff. This Python example submits and polls.

5 min readSume
All posts

fal's H3 Max Lip Sync page says processing takes approximately one minute for 1080p. Sume does not publish a time, and a queued job is normal because concurrency limits apply when workers pick it up, so write your client to poll with exponential backoff and stop on a terminal status. The Python below submits a lip-sync job and waits for the result.

What the vendor page says

The fal page describes one request in, one talking video back, with about a minute of processing at 1080p and output length equal to the audio length (5 to 15 seconds). That is the provider's estimate for the model, not a guarantee for any queue in front of it.

Treat it as the order of magnitude. Plan for tens of seconds to minutes, never milliseconds.

Submit and poll

Sume's job states are queued, processing, completed, failed and canceled. The last three are terminal. Send an Idempotency-Key so a retry after a network timeout cannot create a second paid job. The audio must be Sume-hosted and duration_seconds must be 5 to 14.8.

import json, os, time, urllib.request

API = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
        "Content-Type": "application/json"}

def call(method, path, body=None, key=None):
    headers = dict(HEAD)
    if key:
        headers["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    req = urllib.request.Request(API + path, data=data, headers=headers, method=method)
    with urllib.request.urlopen(req, timeout=60) as resp:
        return json.load(resp)

def main():
    body = {"image_url": os.environ["IMAGE_URL"], "audio_url": os.environ["AUDIO_URL"],
            "duration_seconds": 5.84, "resolution": "1080p"}
    sub = call("POST", "/v1/minimax/h3-max/lip-sync", body, "lipsync-demo-001")
    job_id = sub["job"]["id"]
    delay = 3
    while True:
        st = call("GET", f"/v1/jobs/{job_id}/status")
        state = (st.get("job") or st).get("status")
        print(state)
        if state in ("completed", "failed", "canceled"):
            break
        time.sleep(delay)
        delay = min(delay * 2, 30)
    if state == "completed":
        print(json.dumps(call("GET", f"/v1/jobs/{job_id}/result"))[:500])

main()

Cost and failure

A 5.84-second clip bills as ceil(5.84) = 6 seconds. At 1080p that is 6 x $0.16 x 1.25 = $1.20. The API reserves it at admit, captures it on completion and refunds it if the job fails.

Do not resubmit just because your process timed out. Poll the same job id, or resend with the same Idempotency-Key.

Timing and price for a 5.84 s clip (read 2026-10-05)
Itemfal listingSume
Typical time at 1080pAbout one minuteNot published; async job
Billed secondsOutput lengthceil(5.84) = 6
1080p price per second$0.16 list$0.16 x 1.25 = $0.20
Clip cost at 1080p6 x $0.16 = $0.966 x $0.20 = $1.20
On failureNot stated on the pageReservation refunded

Sources

Related posts

More in Developers

All Developers posts

Written by Sume