LTX-2 Dub-It: lip-matched rephrasing vs Sume's dubbing steps

LTX-2's Dub-It pipeline rephrases speech while matching speaker and lips. What the README lists, what it omits, and what Sume's dubbing steps do and do not do.

5 min readSume
All posts

Dub-It is a pipeline in Lightricks' open LTX-2 repository that rephrases speech in a clip while matching the speaker's identity and lip movements, using a single IC-LoRA in two stages on the distilled transformer. Sume does not offer that: its docs describe dubbing as separate steps (transcribe, synthesize, join, with translation left to you), and say video models do not lip-sync to generated speech or to a later voice-over.

Dub-It facts are from the LTX-2 README and repository; the model card for LTX-2.5 gives the checkpoint facts. All were read on 2026-10-03.

What does Dub-It do, exactly?

The README describes DubItPipeline as rephrasing while matching speaker identity and lip movements. It runs as a single IC-LoRA across two stages and needs the distilled transformer, not a separate full model. The audio sync comes from the model's joint audio-video generation, rather than from a separate lip-sync pass.

Three limits are worth stating. The README does not list supported languages, so do not read "rephrasing" as translation into any language you pick. It gives no clip-length cap and no GPU figure for this pipeline. And the LTX-2.5 card puts the weights under a community licence that is free below $10 million in annual revenue and needs a paid agreement above it, measured across the whole entity.

Dub-It against Sume's documented dubbing steps, read 2026-10-03.
QuestionLTX-2 Dub-ItSume docs
One model changes the speech and the mouth?Yes, per the READMENo; steps are separate
Speaker identity kept?README says it matches speaker identityNot described in the pages read
Lip movement matched?Yes, per the READMENot for video models; Avatar videos lip-sync a still to audio
Languages listed?Not statedNot stated for TTS in the pages read
You run the GPU?YesNo, hosted jobs

What does Sume do for a dub today?

Sume's Models overview says every on-camera speaking shot is an Avatar job with an accepted still and TTS, and that video models do not lip-sync to generated TTS or to a later voice-over. So there is no step in the docs that rewrites the speech inside an existing filmed clip and re-animates the mouth.

What you can assemble from the docs is audio-only dubbing. Video inspect can transcribe at $0.01 per audio minute. Audio detach extracts the track as wav or mp3 for $0.01 per job. tts_create is a paid tool in the hosted MCP inventory, and Timeline audio joins up to 20 parts with no re-synthesis. For a talking presenter built from a still, Avatar videos take an audio_url and a measured duration_seconds. Our dubbing pipeline walkthrough shows the sequence.

What is the real difference in the result?

With separate steps the picture is untouched and only the sound changes, so the mouth keeps moving to the original words. That is acceptable for voice-over on B-roll, product shots or a presenter off camera. On a close-up of a face, a mismatch is visible, and that is the case a lip-matched pipeline targets.

The reverse also holds. A pipeline that regenerates the clip changes the picture, not just the speech, so you need to check the output for drift in face, light and background before you ship it. Neither approach removes the need to disclose synthetic audio where a platform requires it.

Should you run Dub-It yourself?

Run it if your clips are close-up speech, you have a GPU that handles the distilled 22B checkpoint, and the licence fits your revenue. Test it on one 10-second clip first and judge the face, the voice and the timing on a screen the size your audience will use.

Use Sume's steps if the speaker is off camera or small in frame, or if you need many clips with a consistent assembly. When a visible lip mismatch matters and you cannot run the model, the honest option is to re-shoot or to use a presenter built from a still, and not to hide the mismatch under a voice-over.

Sources

Related posts

More in Models

All Models posts

Written by Sume