Lip Forcing video lip-sync: 37 GB GPU or a hosted API?

Lip Forcing is an open video-to-video lip-sync model that needs about 37 GB of GPU memory. What it does, what it asks of you, and the hosted alternatives.

5 min readSume
All posts

Run Lip Forcing yourself only if you already have a GPU with roughly 37 GB of free memory and a video you must keep, because it re-syncs the mouth in an existing video to new audio. If you have a still image and an audio file instead, a hosted lip-sync model is much less work. The model page lists Apache-2.0, a 2 denoising-step-per-chunk student distilled from a 14B teacher, and a peak of about 37 GB of GPU memory, rising to about 50 GB if the text encoder runs at request time (page read 2026-10-04).

The distinction that matters is the input. Video-to-video lip-sync starts from footage. Still-plus-audio lip-sync starts from a photo.

What does the Lip Forcing page say?

The Hugging Face page links an arXiv paper, 2606.11180, and describes real-time-oriented video-to-video lip-sync built by distilling a 14B teacher into a few-step student that works in chunks. It lists a 512 by 512 face alignment step. Those are the vendor's own figures, and we have not run the model, so treat the speed claims as unverified.

Lip Forcing, as listed on its model page, read 2026-10-04
ItemListed value
TaskVideo-to-video lip-sync to new audio
LicenceApache-2.0
Teacher size14B parameters
Denoising steps per chunk2
Peak GPU memoryAbout 37 GB, about 50 GB with runtime text encoding
Face alignment512 by 512

What does Sume offer instead?

Sume's docs do not list a video-to-video lip-sync route. What they do list is generation from a still. Avatar 1.0 generates a talking video from a photo, prompt or props plus a script. The model catalogue also routes lip sync from a still and an audio file: veed/fabric-1.0 and minimax/h3-max/lip-sync, which is covered in the lip-sync options post and the cost comparison.

The Video router is for generation and editing, and the docs do not describe it as syncing an existing clip to a later voice-over.

Which one should you pick?

If the footage is not essential, rebuilding the scene from a still is often the cheaper way to get a lip-synced result, and it avoids the hardware question entirely.

  • You have footage of a person and need new words in their mouth: video-to-video, such as Lip Forcing on your own GPU.
  • You have one photo and a script or recording: a hosted still-plus-audio or avatar route.
  • You have no GPU: do not start with the 37 GB model.
  • You need a hosted API call with a price per second: use the hosted routes and read the amount from the catalogue.

What does 37 GB mean in practice?

A GPU with roughly 37 GB free rules out most consumer cards, so you are looking at a data-centre class card or a rented one. Renting has its own costs, from start-up time to idle billing, and a model that needs an extra encoder at request time can climb toward the 50 GB figure on the page. Factor in your volume before you decide: for a few clips a month, a hosted route you pay per second is usually simpler than keeping a large GPU warm.

If you do self-host, test with your real footage first. Face angle, occlusion and lighting are where lip-sync models differ, and a demo reel will not show you that.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume