Wav2Lip is research-only: what to use for commercial lip sync

Wav2Lip's README limits results to research, academic or personal use. For client or ad work, here is the still-plus-audio lip sync Sume offers, with costs.

5 min readSume
All posts

Wav2Lip is a free research model, and its README says results from the open-source code or its demo site "should only be used for research/academic/personal purposes only". If you are making a paid ad, a client video or a product page, that sentence is the reason to look at a hosted route instead. The README itself points commercial users to Sync Labs, a hosted service, and says commercial use needs contact with the developers.

This post reads the Wav2Lip README on 2026-10-04 and compares it with what Sume documents for lip sync. It is not legal advice: ask a lawyer to read the license for your case.

What does Wav2Lip need as input?

Wav2Lip takes a video of a face plus an audio file, and it changes the mouth in the video to match the audio. The README advises that 720p videos often give better results than 1080p, so you may need to resize your footage first.

  • Input: an existing video of a person, plus an audio file.
  • Setup: you install and run it yourself, on your own GPU.
  • Terms: research, academic or personal use only, per the README.

What does Sume take instead?

Sume's lip-sync routes start from a picture, not a video. POST /v1/veed/fabric-1.0 and POST /v1/minimax/h3-max/lip-sync take one still (a public HTTPS image_url, or a ready avatar by avatar_handle) and an audio_url on a Sume media host, and return a talking clip whose length follows the audio. Fabric accepts audio up to 10 MB; H3 Max accepts 5 to 14.8 seconds. Both run as jobs you poll or receive by webhook.

Wav2Lip README and Sume lip-sync routes (read 2026-10-04)
QuestionWav2LipSume Fabric / H3 Max
Visual inputAn existing video of a faceOne still image or a ready avatar
Audio inputA local audio fileSume-hosted audio_url, 10 MB cap
Where it runsYour GPUSume job, polled or webhook
PriceYour hardware5 s at 720p Fabric $0.94; 5 s at 768p H3 Max $0.50
Stated termsResearch, academic, personalPaid API under Sume terms

Which should you pick?

If you already have footage of a real person and need their mouth re-timed, Wav2Lip is the research baseline, but its stated terms keep it out of paid work, and Sume does not document a video-in route. If you can start from a still, such as a presenter photo or a ready avatar, the Sume route gives you a finished MP4 without a GPU. Make the audio first with Sume text to speech, check its length against the 5 to 14.8 second window for H3 Max, and use Fabric outside that window.

For a script of up to 60 seconds with no audio prepared, skip the two-step route: POST /v1/avatar-1.0/talking-video writes the voice and the face together, from $0.184 per second at the standard tier. Get consent from whoever is in the photo before you animate it.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume