Lip sync AI API options: photo or video input, limits, price

Lip sync APIs take a face and audio and return a talking clip. Some animate a photo, others redub an existing video. Inputs, limits and price units.

5 min readSume
All posts

A lip sync AI API takes a face and an audio track and returns a video whose mouth moves with the audio. The options split by what the face is: a still photo or avatar that gets animated (Creatify Aurora, HeyGen Audio to Video, and Sume's VEED Fabric 1.0 and MiniMax H3 Max Lip Sync), or an existing video whose mouth is redrawn to new audio (HeyGen Lipsync, Kling Lip Sync). Model hosts such as fal and Replicate list many lip sync models, each with its own price.

Vendor facts come from each vendor's own pages, read on 2026-09-28. Sume's come from the Models overview and the request schemas in the Sume API reference.

Which lip sync APIs take a photo, and which take a video?

From Creatify's Aurora, HeyGen's Audio to Video, Lipsync and API pricing, Kling's Lip Sync, face recognition and pricing pages, and Sume's Models overview and API reference, read 2026-09-28.
APIFace inputAudio inputPrice
Creatify AuroraOne photo: png, jpg, jpeg or webpAn mp3 or wav URL, up to 5 minutes1 credit per second, or 0.5 with aurora_v1_fast, from an API plan
HeyGen Audio to VideoAn avatar you own, or any single-person imageA public HTTPS URL, or an uploaded MP3 or WAV up to 32 MB; up to 30 minutes per requestRuns on POST /v3/videos, the avatar video endpoint, which HeyGen prices per minute by engine and avatar type
HeyGen LipsyncAn existing videoReplacement audioListed as Lipsync Audio mode: $0.84 per minute (Speed) or $1.74 (Precision)
Kling Lip SyncAn existing video: mp4 or mov, up to 100 MB, 2–60 seconds, 720p or 1080p; one face per taskmp3, wav, m4a or aac, up to 5 MB, 2–60 seconds$0.07 per 5 seconds, plus $0.007 per face-recognition call
Sume MiniMax H3 Max Lip SyncA public HTTPS still, or a ready avatarAudio on Sume's media host, 5–14.8 seconds$0.10 per audio second (768p)
Sume VEED Fabric 1.0A public HTTPS still, or a ready avatarAudio on Sume's media host, up to 10 MB and 300 seconds$0.1875 per audio second (720p)

What about fal and Replicate?

Both are model hosts that list lip sync models from several model makers. fal lists them under a lipsync category and, per its pricing docs, bills each model by its own unit, only for successful outputs. Replicate's lipsync collection covers models that sync lips in videos or images to new audio; its pricing page says some models bill by hardware time and others by input and output. Read the model's own page for its inputs and price.

How do I choose a lip sync API?

  • Start from what you have. A still photo or an avatar needs a photo-driven API; a finished video you want to redub needs HeyGen Lipsync, Kling, or a hosted video-to-video model. Sume's two endpoints take a still or an avatar, not existing footage.
  • Check where the audio must live. Sume refuses audio that is not on its media host, so the usual source is a Sume TTS clip; there is no public upload for a local file. HeyGen takes a public URL or an upload, and Kling a URL or Base64.
  • Check the length caps against your longest line: 2–60 seconds on Kling, 5–14.8 on H3 Max, up to 300 on Fabric, 5 minutes on Aurora, 30 minutes on HeyGen Audio to Video.
  • Compare the unit: per second, per 5 seconds, per minute or per credit. Sume's rates are before a 5.5% agent fee by default.

Can a video model lip-sync to my voice-over instead?

Not on Sume. Its models guide says video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is built from a still plus audio on Fabric. The request, limits and a worked price are in Lip sync API: image and audio to a talking clip, and lip sync vs dubbing explains when you need which.

Sources

Related posts

More in Models

All Models posts

Written by Sume