NVIDIA LipSync NIM: de/es/fr models vs Sume lip sync

NVIDIA's LipSync NIM re-times mouths in existing video to new audio. Sume lip sync starts from a still and an audio file, so existing footage is not re-synced.

5 min readSume
All posts

Can Sume do what NVIDIA's LipSync NIM does?

No. NVIDIA's LipSync NIM takes a finished video plus a new speech track and changes the mouth movement in that video. Sume does not offer that: its lip sync starts from a still image (or a saved avatar) and an audio file, and produces a new talking clip.

That difference decides whether the tool belongs in your dubbing workflow. If you already have a filmed presenter and want their mouth to match a German track, you need a video-to-video lip-sync service. If you are making the speaking clip from a picture and a script in the first place, Sume's route fits, and the language of the clip is just the language of the audio you feed it.

What did NVIDIA announce for LipSync NIM at IBC 2026?

NVIDIA's IBC 2026 post (IBC ran September 11-14 in Amsterdam) describes the LipSync NIM microservice as transforming mouth movement in an input video to match a target audio track while preserving head pose, blinking and body movement. The newest release is described as better at partially obscured faces. NDI is using it for real-time translation and dubbing inside broadcast workflows.

NVIDIA's documentation for release 1.3.0 lists a language-agnostic generic model plus fine-tuned models for German (de), Spanish (es) and French (fr). The overview says it accepts a video, a speech file and, optionally, per-frame speaker bounding boxes for multi-speaker scenes, and offers a streaming mode (recommended) and a transactional mode that needs the whole file first.

The support matrix is also a reminder that this is self-hosted GPU software: it lists T4, A2, A10, A16, A40, L4, L40, L40s and B40 server cards plus RTX 4090, 5080 and 5090, and states that GPUs without NVENC/NVDEC, including A100, H100 and B100 products, are not supported. The page does not state video resolution or duration limits, so do not assume any.

What does Sume ship for talking video?

Per the models overview, every on-camera speaking shot is a Fabric clip made from an accepted still plus TTS audio: you send audio_url, a measured duration_seconds and exactly one visual source (image_url or avatar_handle). The same page says video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is never a video-model clip with narration laid underneath.

For script-driven avatar clips, the avatar video docs accept scripts whose estimated length is 4 to 60 seconds. Longer scripts have to be shortened or split into several jobs.

There is no endpoint that accepts an existing MP4 and a replacement audio track and returns the same video with new mouth movement. If a tool or a post claims otherwise for Sume, check it against the docs before you plan around it.

NVIDIA LipSync NIM and Sume lip sync, read 2026-10-02
QuestionNVIDIA LipSync NIM 1.3.0Sume
InputExisting video plus speech audioA still or saved avatar plus audio (or a script)
Changes mouth in existing footageYesNo
Language handlingGeneric model; fine-tuned de, es, frFollows the audio you supply; set language on TTS
Where it runsYour own NVENC-capable NVIDIA GPUsSume API jobs
Length limits statedNot stated on the support matrixAvatar scripts 4-60 seconds

How do you localize a speaking clip on Sume instead?

Treat each language as a new take, not a patch on the old one. Transcribe the source speech with Sume STT if you need the words (detach the audio first with format: wav, channels: mono, sample_rate: 16000), translate the text outside Sume or with an agent step you review, generate the target-language voice with TTS, then render the speaking clip from the same still.

Set language on the TTS request for every non-English transcript. The OpenAPI description says an omitted value defaults to English at the provider, and a voice whose language does not match gets a warning you have to confirm. Our dubbing pipeline walk-through shows the chain end to end, and lip sync vs dubbing explains when each is worth paying for.

  • Same face, new language: re-render from the same still with the new audio.
  • Original footage must stay as filmed: dub the audio track only and accept visible mismatch, or use a video-to-video service.
  • Timing: a translated line is rarely the same length, so cut by sentence segments (see keeping translated lines on time).

When should you choose NVIDIA's NIM instead?

Pick LipSync NIM when the footage is the asset: interviews, news, sports or a recorded executive whose mouth must match a German, Spanish or French track, and you can run NVENC-capable GPUs or reach them through a partner. Pick Sume when the speaking clip is generated from a still, you want a hosted job with a webhook, and a 4-60 second clip per job fits the content. Many teams will use both: generated shorts on Sume, filmed hero footage through a broadcast-grade lip-sync stack.

Whichever you use, a synthetic voice or altered mouth movement can trigger disclosure rules on some platforms. Check the platform and jurisdiction before publishing.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume