Lip Forcing video lip-sync: 37 GB GPU or a hosted API?
Lip Forcing is an open video-to-video lip-sync model that needs about 37 GB of GPU memory. What it does, what it asks of you, and the hosted alternatives.

Run Lip Forcing yourself only if you already have a GPU with roughly 37 GB of free memory and a video you must keep, because it re-syncs the mouth in an existing video to new audio. If you have a still image and an audio file instead, a hosted lip-sync model is much less work. The model page lists Apache-2.0, a 2 denoising-step-per-chunk student distilled from a 14B teacher, and a peak of about 37 GB of GPU memory, rising to about 50 GB if the text encoder runs at request time (page read 2026-10-04).
The distinction that matters is the input. Video-to-video lip-sync starts from footage. Still-plus-audio lip-sync starts from a photo.
What does the Lip Forcing page say?
The Hugging Face page links an arXiv paper, 2606.11180, and describes real-time-oriented video-to-video lip-sync built by distilling a 14B teacher into a few-step student that works in chunks. It lists a 512 by 512 face alignment step. Those are the vendor's own figures, and we have not run the model, so treat the speed claims as unverified.
| Item | Listed value |
|---|---|
| Task | Video-to-video lip-sync to new audio |
| Licence | Apache-2.0 |
| Teacher size | 14B parameters |
| Denoising steps per chunk | 2 |
| Peak GPU memory | About 37 GB, about 50 GB with runtime text encoding |
| Face alignment | 512 by 512 |
What does Sume offer instead?
Sume's docs do not list a video-to-video lip-sync route. What they do list is generation from a still. Avatar 1.0 generates a talking video from a photo, prompt or props plus a script. The model catalogue also routes lip sync from a still and an audio file: veed/fabric-1.0 and minimax/h3-max/lip-sync, which is covered in the lip-sync options post and the cost comparison.
The Video router is for generation and editing, and the docs do not describe it as syncing an existing clip to a later voice-over.
Which one should you pick?
If the footage is not essential, rebuilding the scene from a still is often the cheaper way to get a lip-synced result, and it avoids the hardware question entirely.
- You have footage of a person and need new words in their mouth: video-to-video, such as Lip Forcing on your own GPU.
- You have one photo and a script or recording: a hosted still-plus-audio or avatar route.
- You have no GPU: do not start with the 37 GB model.
- You need a hosted API call with a price per second: use the hosted routes and read the amount from the catalogue.
What does 37 GB mean in practice?
A GPU with roughly 37 GB free rules out most consumer cards, so you are looking at a data-centre class card or a rented one. Renting has its own costs, from start-up time to idle billing, and a model that needs an extra encoder at request time can climb toward the 50 GB figure on the page. Factor in your volume before you decide: for a few clips a month, a hosted route you pay per second is usually simpler than keeping a large GPU warm.
If you do self-host, test with your real footage first. Face angle, occlusion and lighting are where lip-sync models differ, and a demo reel will not show you that.
Sources
Related posts
More in Comparisons
- Live avatar agent or rendered video: D-ID vs Sume
D-ID V4 Expressive Visual Agents are live, LLM-connected avatars. Sume avatar video is rendered from a script in 4 to 60 s. Which one fits which job.
- How long can an AI song be? Suno, ElevenLabs and Lyria compared
ElevenLabs Music runs 3 s to 5 minutes; Suno v6 is reported at up to 8 minutes; Lyria 3.5 makes a couple of minutes. Join tracks with a $0.01 concat.
- LTX-2.5 license ($10M ARR line) vs hosted video ids
LTX-2.x Community license is free under $10M ARR and paid above it. Here is what changes if you call a hosted Sume video id instead of running weights.
- LTX-2.5 Fast 20 s vs Pro 10 s: long takes on Sume
LTX-2.5 Fast runs up to 20 s at 720p and 1080p, Pro 6 to 10 s. Sume lists no LTX id, but wan-3.0 and seedance-2.5 take single clips up to 30 s.
Written by Sume