Wav2Lip is research-only: what to use for commercial lip sync
Wav2Lip's README limits results to research, academic or personal use. For client or ad work, here is the still-plus-audio lip sync Sume offers, with costs.

Wav2Lip is a free research model, and its README says results from the open-source code or its demo site "should only be used for research/academic/personal purposes only". If you are making a paid ad, a client video or a product page, that sentence is the reason to look at a hosted route instead. The README itself points commercial users to Sync Labs, a hosted service, and says commercial use needs contact with the developers.
This post reads the Wav2Lip README on 2026-10-04 and compares it with what Sume documents for lip sync. It is not legal advice: ask a lawyer to read the license for your case.
What does Wav2Lip need as input?
Wav2Lip takes a video of a face plus an audio file, and it changes the mouth in the video to match the audio. The README advises that 720p videos often give better results than 1080p, so you may need to resize your footage first.
- Input: an existing video of a person, plus an audio file.
- Setup: you install and run it yourself, on your own GPU.
- Terms: research, academic or personal use only, per the README.
What does Sume take instead?
Sume's lip-sync routes start from a picture, not a video. POST /v1/veed/fabric-1.0 and POST /v1/minimax/h3-max/lip-sync take one still (a public HTTPS image_url, or a ready avatar by avatar_handle) and an audio_url on a Sume media host, and return a talking clip whose length follows the audio. Fabric accepts audio up to 10 MB; H3 Max accepts 5 to 14.8 seconds. Both run as jobs you poll or receive by webhook.
| Question | Wav2Lip | Sume Fabric / H3 Max |
|---|---|---|
| Visual input | An existing video of a face | One still image or a ready avatar |
| Audio input | A local audio file | Sume-hosted audio_url, 10 MB cap |
| Where it runs | Your GPU | Sume job, polled or webhook |
| Price | Your hardware | 5 s at 720p Fabric $0.94; 5 s at 768p H3 Max $0.50 |
| Stated terms | Research, academic, personal | Paid API under Sume terms |
Which should you pick?
If you already have footage of a real person and need their mouth re-timed, Wav2Lip is the research baseline, but its stated terms keep it out of paid work, and Sume does not document a video-in route. If you can start from a still, such as a presenter photo or a ready avatar, the Sume route gives you a finished MP4 without a GPU. Make the audio first with Sume text to speech, check its length against the 5 to 14.8 second window for H3 Max, and use Fabric outside that window.
For a script of up to 60 seconds with no audio prepared, skip the two-step route: POST /v1/avatar-1.0/talking-video writes the voice and the face together, from $0.184 per second at the standard tier. Get consent from whoever is in the photo before you animate it.
Sources
Related posts
More in Sume Avatar 1.0
- Alt text for an avatar video poster: WCAG 1.1.1 in practice
A poster image from a Sume avatar preview still needs a text alternative under WCAG 1.1.1, and the video needs descriptive identification. What to write.
- Silent avatar video clip: what WCAG 1.2.1 asks for
A silent beat or silent clip in an avatar video is prerecorded video-only content. Here is what WCAG 1.2.1 wants and how to supply it from Sume.
- WCAG 1.2.3 for an AI avatar video: is the script enough?
A talking-head avatar video already has its words in a script. Here is when that text meets WCAG 1.2.3 and what to add if the picture carries more.
- Zoom Clips custom avatar: 18 minutes a month vs per-second billing
Zoom's Custom Avatar add-on lists 5 avatars and 18 minutes of AI video a month. See how that compares with paying Sume per second, with worked costs.
Written by Sume