Can you lip-sync an existing video to new audio with Sume?
Sume's docs list no video-to-video lip-sync route. Here are the three documented ways to get a speaking face: from a still, face swap, or a new avatar video.

No: the Sume docs we checked list no route that takes an existing video and re-syncs its mouth to a new audio track. What they do list are three other ways to end up with a speaking face. You can generate a talking video from a still image, swap a face into a clip, or generate a fresh avatar video from a script.
That is a real limit, and it is worth knowing before you plan a dubbing workflow around it. Video generation models in the router are for creating or editing footage, and the docs do not describe them as lip-syncing to a later voice-over.
What are the three documented routes?
Each starts from a different input, so pick by what you already have.
| You have | Route | What you get |
|---|---|---|
| A photo and a script | Avatar 1.0 generate or talking-video | A new talking video; script must estimate at 4 to 60 seconds |
| A still and a recording | Fabric 1.0 or H3 Max Lip Sync from the model catalogue | A video of that still speaking the audio |
| A clip with a person and a source face | Avatar face swap (Beta) | The clip with the face replaced; needs about 4 to 15 seconds with audio |
Which one is closest to dubbing?
None is true dubbing, where the original person's mouth is re-timed to a translated voice. The closest for a spokesperson workflow is to rebuild the spokesperson as an avatar and feed it the translated script. That changes the footage rather than preserving it, so it only works when the person on screen can be a generated one.
If you must keep real footage and change the words, you need a video-to-video lip-sync model outside Sume, and you should check its hardware needs first. Open models of that kind exist, with a typical cost in GPU memory, as described in the Lip Forcing comparison.
What can Sume still do for a translated clip?
For the subtitle path, see the lip-sync options overview for what each still-based route costs, and the face-swap post for what the Beta needs as input. Read the Avatar 1.0 docs for the current script, ratio and quality options before you submit.
- Burn translated subtitles onto the original video with the caption endpoint, no lip movement needed.
- Detach the original audio and split or join audio parts when you build a new track.
- Generate a new talking-head version from a still when brand rules allow an avatar.
What should you ask before choosing a route?
Ask first whether the person on screen must be the real person. If yes, an existing-footage tool is needed and you should look outside Sume for the mouth edit, then use Sume for captions and audio steps. If an avatar is acceptable, ask whether you have the photo rights and whether the script fits the 4 to 60 second window. Those two questions decide the route in most cases.
If you are not sure, run a three-second test on the cheapest route and watch it before committing a whole campaign.
Sources
Related posts
More in Use cases
- Live-commerce host video by API: instruction, input, schema, webhook
How a Sume Format run for a live-commerce style host video is shaped: instruction for decisions, input for data, an output schema, a spend cap and a webhook.
- Lofi study stream: a 30-second track looped as a Sume soundtrack
A lo-fi study stream needs a bed that repeats cleanly. Generate a 30-second track with Sume's router and loop it under a video with soundtrack.loop.
- Lyria 3.5 screens prompts: brand slogans and lyrics in an ad song
Google says every Lyria prompt is safety-screened and it cannot output copyrighted lyrics or named artist voices. How to write a brand jingle brief that passes.
- Make 3-minute hold music for a voice agent
ElevenLabs agents can play hold audio up to 180 s and 40 MB. Generate a short loop with Sume music generation and join it into 180 s with timeline audio.
Written by Sume