Lip sync AI API options: photo or video input, limits, price
Lip sync APIs take a face and audio and return a talking clip. Some animate a photo, others redub an existing video. Inputs, limits and price units.

A lip sync AI API takes a face and an audio track and returns a video whose mouth moves with the audio. The options split by what the face is: a still photo or avatar that gets animated (Creatify Aurora, HeyGen Audio to Video, and Sume's VEED Fabric 1.0 and MiniMax H3 Max Lip Sync), or an existing video whose mouth is redrawn to new audio (HeyGen Lipsync, Kling Lip Sync). Model hosts such as fal and Replicate list many lip sync models, each with its own price.
Vendor facts come from each vendor's own pages, read on 2026-09-28. Sume's come from the Models overview and the request schemas in the Sume API reference.
Which lip sync APIs take a photo, and which take a video?
| API | Face input | Audio input | Price |
|---|---|---|---|
| Creatify Aurora | One photo: png, jpg, jpeg or webp | An mp3 or wav URL, up to 5 minutes | 1 credit per second, or 0.5 with aurora_v1_fast, from an API plan |
| HeyGen Audio to Video | An avatar you own, or any single-person image | A public HTTPS URL, or an uploaded MP3 or WAV up to 32 MB; up to 30 minutes per request | Runs on POST /v3/videos, the avatar video endpoint, which HeyGen prices per minute by engine and avatar type |
| HeyGen Lipsync | An existing video | Replacement audio | Listed as Lipsync Audio mode: $0.84 per minute (Speed) or $1.74 (Precision) |
| Kling Lip Sync | An existing video: mp4 or mov, up to 100 MB, 2–60 seconds, 720p or 1080p; one face per task | mp3, wav, m4a or aac, up to 5 MB, 2–60 seconds | $0.07 per 5 seconds, plus $0.007 per face-recognition call |
| Sume MiniMax H3 Max Lip Sync | A public HTTPS still, or a ready avatar | Audio on Sume's media host, 5–14.8 seconds | $0.10 per audio second (768p) |
| Sume VEED Fabric 1.0 | A public HTTPS still, or a ready avatar | Audio on Sume's media host, up to 10 MB and 300 seconds | $0.1875 per audio second (720p) |
What about fal and Replicate?
Both are model hosts that list lip sync models from several model makers. fal lists them under a lipsync category and, per its pricing docs, bills each model by its own unit, only for successful outputs. Replicate's lipsync collection covers models that sync lips in videos or images to new audio; its pricing page says some models bill by hardware time and others by input and output. Read the model's own page for its inputs and price.
How do I choose a lip sync API?
- Start from what you have. A still photo or an avatar needs a photo-driven API; a finished video you want to redub needs HeyGen Lipsync, Kling, or a hosted video-to-video model. Sume's two endpoints take a still or an avatar, not existing footage.
- Check where the audio must live. Sume refuses audio that is not on its media host, so the usual source is a Sume TTS clip; there is no public upload for a local file. HeyGen takes a public URL or an upload, and Kling a URL or Base64.
- Check the length caps against your longest line: 2–60 seconds on Kling, 5–14.8 on H3 Max, up to 300 on Fabric, 5 minutes on Aurora, 30 minutes on HeyGen Audio to Video.
- Compare the unit: per second, per 5 seconds, per minute or per credit. Sume's rates are before a 5.5% agent fee by default.
Can a video model lip-sync to my voice-over instead?
Not on Sume. Its models guide says video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is built from a still plus audio on Fabric. The request, limits and a worked price are in Lip sync API: image and audio to a talking clip, and lip sync vs dubbing explains when you need which.
Sources
- Creatify: create an Aurora task (read 2026-09-28)
- HeyGen Audio to Video (read 2026-09-28)
- HeyGen Lipsync (read 2026-09-28)
- HeyGen API pricing explained (read 2026-09-28)
- Kling Lip Sync API (read 2026-09-28)
- Kling Lip Sync face recognition (read 2026-09-28)
- Kling API video pricing (read 2026-09-28)
- fal lipsync models (read 2026-09-28)
- fal Model APIs pricing (read 2026-09-28)
- Replicate lipsync collection (read 2026-09-28)
- Replicate pricing (read 2026-09-28)
- Models overview
- Sume API reference
- API reference
- Sume API pricing
Related posts
More in Models
- Lyria MCP: generate music with Lyria 3.5 from Claude
Sume's hosted MCP server has a music_create tool that runs on Google Lyria 3.5 today: a prompt in, a job id back, an audio file on media.sume.com.
- Multi-shot video generation: what it is and how to do it
Multi-shot video generation returns one clip with several cuts. Which APIs offer it (Kling 3.0, Runway) and how to build it shot by shot on Sume.
- Nano Banana MCP server for Claude: how to add one
To use Nano Banana in Claude, connect an MCP server that runs it. Which servers list it, how to add Sume's, and what a call looks like and costs.
- One API for Kling, Seedance and other video models, one bill
Yes: Sume's POST /v1/videos calls Kling, Seedance and other video models with one API key and one bill. Veo is not in its catalog. Rates and a sample call.
Written by Sume