LTX-2 Dub-It: lip-matched rephrasing vs Sume's dubbing steps
LTX-2's Dub-It pipeline rephrases speech while matching speaker and lips. What the README lists, what it omits, and what Sume's dubbing steps do and do not do.

Dub-It is a pipeline in Lightricks' open LTX-2 repository that rephrases speech in a clip while matching the speaker's identity and lip movements, using a single IC-LoRA in two stages on the distilled transformer. Sume does not offer that: its docs describe dubbing as separate steps (transcribe, synthesize, join, with translation left to you), and say video models do not lip-sync to generated speech or to a later voice-over.
Dub-It facts are from the LTX-2 README and repository; the model card for LTX-2.5 gives the checkpoint facts. All were read on 2026-10-03.
What does Dub-It do, exactly?
The README describes DubItPipeline as rephrasing while matching speaker identity and lip movements. It runs as a single IC-LoRA across two stages and needs the distilled transformer, not a separate full model. The audio sync comes from the model's joint audio-video generation, rather than from a separate lip-sync pass.
Three limits are worth stating. The README does not list supported languages, so do not read "rephrasing" as translation into any language you pick. It gives no clip-length cap and no GPU figure for this pipeline. And the LTX-2.5 card puts the weights under a community licence that is free below $10 million in annual revenue and needs a paid agreement above it, measured across the whole entity.
| Question | LTX-2 Dub-It | Sume docs |
|---|---|---|
| One model changes the speech and the mouth? | Yes, per the README | No; steps are separate |
| Speaker identity kept? | README says it matches speaker identity | Not described in the pages read |
| Lip movement matched? | Yes, per the README | Not for video models; Avatar videos lip-sync a still to audio |
| Languages listed? | Not stated | Not stated for TTS in the pages read |
| You run the GPU? | Yes | No, hosted jobs |
What does Sume do for a dub today?
Sume's Models overview says every on-camera speaking shot is an Avatar job with an accepted still and TTS, and that video models do not lip-sync to generated TTS or to a later voice-over. So there is no step in the docs that rewrites the speech inside an existing filmed clip and re-animates the mouth.
What you can assemble from the docs is audio-only dubbing. Video inspect can transcribe at $0.01 per audio minute. Audio detach extracts the track as wav or mp3 for $0.01 per job. tts_create is a paid tool in the hosted MCP inventory, and Timeline audio joins up to 20 parts with no re-synthesis. For a talking presenter built from a still, Avatar videos take an audio_url and a measured duration_seconds. Our dubbing pipeline walkthrough shows the sequence.
What is the real difference in the result?
With separate steps the picture is untouched and only the sound changes, so the mouth keeps moving to the original words. That is acceptable for voice-over on B-roll, product shots or a presenter off camera. On a close-up of a face, a mismatch is visible, and that is the case a lip-matched pipeline targets.
The reverse also holds. A pipeline that regenerates the clip changes the picture, not just the speech, so you need to check the output for drift in face, light and background before you ship it. Neither approach removes the need to disclose synthetic audio where a platform requires it.
Should you run Dub-It yourself?
Run it if your clips are close-up speech, you have a GPU that handles the distilled 22B checkpoint, and the licence fits your revenue. Test it on one 10-second clip first and judge the face, the voice and the timing on a screen the size your audience will use.
Use Sume's steps if the speaker is off camera or small in frame, or if you need many clips with a consistent assembly. When a visible lip mismatch matters and you cannot run the model, the honest option is to re-shoot or to use a presenter built from a still, and not to hide the mismatch under a voice-over.
Sources
Related posts
More in Models
- Luma Ray 3.2 edit controls: pose, depth, normals, and Sume's edit
Luma's Ray 3.2 edit takes auto_controls, nine strength presets, or per-signal pose, depth, normals, trajectory and face controls. Sume's edit is prompt-only.
- MiniMax-H3 Turbo LoRA: 4 to 8 steps, Apache 2.0, base licence
The MiniMax-H3 Turbo LoRA cuts sampling to 4 to 8 steps and lists Apache 2.0, but the 33B base keeps its own terms. What it needs and the hosted route.
- MiniMax-Music3 weights: 5-minute songs, 8 GB, vs Sume Music Router
MiniMax-Music3 is open-weights: 5-minute 32 kHz songs, 8 GB with streaming, a $20M licence rule. What it needs, and what Sume's Music Router does instead.
- Nano Banana Pro interleaved text and images vs Sume image output
Google documents Nano Banana Pro returning text blocks with illustrations in one answer. Sume's Image API documents image results only; here is the workaround.
Written by Sume