LatentSync 1.6 needs 18 GB VRAM: a hosted lip sync from a still

LatentSync 1.6 is a 512 px video-plus-audio lip-sync model that wants 18 GB of GPU memory. No GPU, only a photo? Here is the Sume still-plus-audio route.

5 min readSume
All posts

LatentSync is ByteDance's Apache-2.0 lip-sync model. Version 1.6, released June 2025 per the README, was trained at 512 x 512 to fix the blur of version 1.5, and the same README lists 18 GB of VRAM as the minimum for inference, against 8 GB for 1.5. If you do not have an 18 GB card, or you do not want to rent one for a few clips, the question becomes what a hosted route can do.

The LatentSync facts come from its README, read 2026-10-04. The Sume facts come from the models page and the routes in the Sume repository.

What LatentSync takes

It takes a video file and an audio track, and re-draws the mouth region so it follows the audio. It exposes inference_steps (20 to 50) and guidance_scale (1.0 to 3.0): more steps improve quality, a higher guidance scale improves lip-sync accuracy. Output is 256 x 256 in 1.5 and 512 x 512 in 1.6.

LatentSync README versions (read 2026-10-04)
VersionInference VRAMOutput face resolution
1.58 GB minimum256 x 256
1.618 GB minimum512 x 512

When a hosted still route fits

LatentSync needs existing footage. If what you have is a portrait, an avatar or a product mascot, there is no video to feed it. Sume's lip-sync routes start from a still instead: POST /v1/veed/fabric-1.0 (audio up to 10 MB, 480p or 720p) or POST /v1/minimax/h3-max/lip-sync (audio of 5 to 14.8 seconds, 480p, 768p or 1080p). You send a public HTTPS image or a ready avatar handle plus audio already hosted on Sume media, then poll the job or take a webhook. Nothing about GPU memory is yours to manage.

Costs are per second of audio. A 5 second Fabric clip at 720p reserves $0.94, and a 5 second H3 Max clip at 768p reserves $0.50. The reservation is captured on completion and refunded on failure.

What you give up

You give up control. There is no guidance_scale to turn, and the route does not re-time a mouth in your own footage. If the shot is a real person's recorded video and only the lips must change, a self-hosted model is the closer fit. If the shot is a still that should speak, the hosted job is simpler. For a full scripted presenter, POST /v1/avatar-1.0/talking-video covers 4 to 60 seconds in one request, so you skip the separate audio step.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume