How long does AI lip sync take? fal says about a minute at 1080p
fal says a 1080p H3 Max lip-sync clip takes about a minute. On Sume the call is an async job, so poll with backoff. This Python example submits and polls.

fal's H3 Max Lip Sync page says processing takes approximately one minute for 1080p. Sume does not publish a time, and a queued job is normal because concurrency limits apply when workers pick it up, so write your client to poll with exponential backoff and stop on a terminal status. The Python below submits a lip-sync job and waits for the result.
What the vendor page says
The fal page describes one request in, one talking video back, with about a minute of processing at 1080p and output length equal to the audio length (5 to 15 seconds). That is the provider's estimate for the model, not a guarantee for any queue in front of it.
Treat it as the order of magnitude. Plan for tens of seconds to minutes, never milliseconds.
Submit and poll
Sume's job states are queued, processing, completed, failed and canceled. The last three are terminal. Send an Idempotency-Key so a retry after a network timeout cannot create a second paid job. The audio must be Sume-hosted and duration_seconds must be 5 to 14.8.
import json, os, time, urllib.request
API = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"}
def call(method, path, body=None, key=None):
headers = dict(HEAD)
if key:
headers["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
req = urllib.request.Request(API + path, data=data, headers=headers, method=method)
with urllib.request.urlopen(req, timeout=60) as resp:
return json.load(resp)
def main():
body = {"image_url": os.environ["IMAGE_URL"], "audio_url": os.environ["AUDIO_URL"],
"duration_seconds": 5.84, "resolution": "1080p"}
sub = call("POST", "/v1/minimax/h3-max/lip-sync", body, "lipsync-demo-001")
job_id = sub["job"]["id"]
delay = 3
while True:
st = call("GET", f"/v1/jobs/{job_id}/status")
state = (st.get("job") or st).get("status")
print(state)
if state in ("completed", "failed", "canceled"):
break
time.sleep(delay)
delay = min(delay * 2, 30)
if state == "completed":
print(json.dumps(call("GET", f"/v1/jobs/{job_id}/result"))[:500])
main()Cost and failure
A 5.84-second clip bills as ceil(5.84) = 6 seconds. At 1080p that is 6 x $0.16 x 1.25 = $1.20. The API reserves it at admit, captures it on completion and refunds it if the job fails.
Do not resubmit just because your process timed out. Poll the same job id, or resend with the same Idempotency-Key.
| Item | fal listing | Sume |
|---|---|---|
| Typical time at 1080p | About one minute | Not published; async job |
| Billed seconds | Output length | ceil(5.84) = 6 |
| 1080p price per second | $0.16 list | $0.16 x 1.25 = $0.20 |
| Clip cost at 1080p | 6 x $0.16 = $0.96 | 6 x $0.20 = $1.20 |
| On failure | Not stated on the page | Reservation refunded |
Sources
Related posts
More in Developers
- How long to wait for a Sume job webhook before you start polling
Job webhooks retry 10 times, 30 seconds apart, with a 10-second timeout each. Start your poll fallback at about 6 minutes, and here is the arithmetic.
- How many jobs_wait calls does a long video job need? 50 s slices
At the default 50 s slice a 10-minute render needs up to 12 jobs_wait calls, at the 55 s cap up to 11. Write the call budget into the agent instruction.
- How to cap a Sume Format run's spend: $500 ceiling, $400 default
Send generation_spend_cap_usd on each Format run. A value above 500 returns 400, a Format with no cap defaults to 400, and a run past its cap fails.
- Idempotency-Key per shot: rerun one failed AI video shot in Python
One key per shot, derived from project, shot number and revision: a retry returns the same Sume job, a changed prompt gets a new key. Python key helper inside.
Written by Sume