LatentSync 1.6 needs 18 GB VRAM: a hosted lip sync from a still
LatentSync 1.6 is a 512 px video-plus-audio lip-sync model that wants 18 GB of GPU memory. No GPU, only a photo? Here is the Sume still-plus-audio route.

LatentSync is ByteDance's Apache-2.0 lip-sync model. Version 1.6, released June 2025 per the README, was trained at 512 x 512 to fix the blur of version 1.5, and the same README lists 18 GB of VRAM as the minimum for inference, against 8 GB for 1.5. If you do not have an 18 GB card, or you do not want to rent one for a few clips, the question becomes what a hosted route can do.
The LatentSync facts come from its README, read 2026-10-04. The Sume facts come from the models page and the routes in the Sume repository.
What LatentSync takes
It takes a video file and an audio track, and re-draws the mouth region so it follows the audio. It exposes inference_steps (20 to 50) and guidance_scale (1.0 to 3.0): more steps improve quality, a higher guidance scale improves lip-sync accuracy. Output is 256 x 256 in 1.5 and 512 x 512 in 1.6.
| Version | Inference VRAM | Output face resolution |
|---|---|---|
| 1.5 | 8 GB minimum | 256 x 256 |
| 1.6 | 18 GB minimum | 512 x 512 |
When a hosted still route fits
LatentSync needs existing footage. If what you have is a portrait, an avatar or a product mascot, there is no video to feed it. Sume's lip-sync routes start from a still instead: POST /v1/veed/fabric-1.0 (audio up to 10 MB, 480p or 720p) or POST /v1/minimax/h3-max/lip-sync (audio of 5 to 14.8 seconds, 480p, 768p or 1080p). You send a public HTTPS image or a ready avatar handle plus audio already hosted on Sume media, then poll the job or take a webhook. Nothing about GPU memory is yours to manage.
Costs are per second of audio. A 5 second Fabric clip at 720p reserves $0.94, and a 5 second H3 Max clip at 768p reserves $0.50. The reservation is captured on completion and refunded on failure.
What you give up
You give up control. There is no guidance_scale to turn, and the route does not re-time a mouth in your own footage. If the shot is a real person's recorded video and only the lips must change, a self-hosted model is the closer fit. If the shot is a still that should speak, the hosted job is simpler. For a full scripted presenter, POST /v1/avatar-1.0/talking-video covers 4 to 60 seconds in one request, so you skip the separate audio step.
Sources
Related posts
More in Sume Avatar 1.0
- LivePortrait needs a driving video, not audio: Kling motion control
LivePortrait animates a portrait from a driving video, not from speech. Sume's Kling 3.0 motion control route does that as a hosted API, per second.
- Nonprofit donor thank-you: three tiered avatar videos, total cost
Three 20-second thank-you videos, one per giving tier, cost $11.04 in total at standard. See how to script them and why you review a preview first.
- Audio2Face-3D blendshapes vs a finished lip-synced MP4 from a photo
NVIDIA Audio2Face-3D turns audio into facial animation data for a 3D rig. If you want a finished MP4 from a photo instead, this is the Sume route and its cost.
- Podcast episode promo: a 20-second avatar host clip, cost by tier
Turn an episode summary into a 20-second vertical promo with an AI avatar host. Standard costs $3.68, plus $4.90, max $11.00. Includes captions and a script.
Written by Sume