LTX-2.5 text encoder: use the bundled Gemma 4 12B, not stock
LTX-2.5 needs its own Gemma 4 12B text encoder; Google's stock Gemma 4 release is not a substitute. What to download and what the 66 GB package holds.

LTX-2.5 requires the Gemma 4 12B text encoder that ships in its own model package. The LTX-2 README states that Google's stock Gemma 4 release is not a substitute, because the encoder is trained for this version. If your run fails or your prompts seem ignored, check you did not swap in a generic Gemma checkpoint.
Why does the encoder matter?
The Hugging Face card says LTX-2.5 uses a custom Gemma 4 12B encoder designed to hold complex prompts together, with several characters and camera moves. The text encoder turns your prompt into the conditioning the 22B transformer follows. A different encoder gives the transformer inputs it was not trained on, so output quality can drop in ways that look like a prompt problem.
What is in the model package?
The README describes a package of roughly 66 GB distributed as separate component files, downloadable with the Hugging Face CLI. LTX-2.5 is the current release and LTX-2.3 remains supported as legacy.
| Component | Note |
|---|---|
| Transformer | 22B; distilled (8-step) or full/dev |
| Text encoder | Bundled Gemma 4 12B |
| Video decoder | Conv VAE (faster) or DiffVAE (higher quality) |
| Upscalers | Separate files in the package |
| Quantized variants | NVFP4 and ComfyUI int8 per the card |
How do I point the pipeline at the right files?
The README's example runs the distilled pipeline with --transformer-path set to the 22B distilled bf16 safetensors under models/ltx-2.5/diffusion_models/, and a frame count of 121. Keep the encoder and the transformer from the same download. Hugging Face may ask you to accept terms and log in first; see the gated download post.
What if I do not want to manage 66 GB of files?
That is a fair reason to use a hosted video route. Sume lists its models at GET /v1/videos/models, and the docs pages I read do not list an LTX id, so confirm before you plan around it. The Video generation docs show how to submit and poll. Hosting removes the encoder question, and it also removes your ability to change the weights.
If you stay local, remember the license: free under $10 million annual revenue, a paid agreement above it.
What are the symptoms of a mismatched encoder?
I have not reproduced a mismatch, so I will not describe the exact failure. The README's warning is the useful fact: a generic Gemma 4 release is not a substitute. If a run loads without an error but follows prompts poorly, compare the file hashes and paths of your encoder and transformer against the package you downloaded.
A clean way to avoid the issue is to download the whole package into one directory with the Hugging Face CLI, point every path at that directory, and avoid mixing in files from other projects or older LTX versions. LTX-2.3 is still supported as legacy, so keep its files in a separate folder.
Sources
Related posts
More in Models
- LTX-2.5 duration predictor: optional, and Sume sets seconds
LTX-2.5 has an optional duration predictor that sets the frame count from the prompt. Sume has no such mode: you send duration in whole seconds.
- LTX-2.5 native multishot: one prompt, or clips joined on Sume
LTX-2.5 adds native multishot: connected scenes in one pass, consistent characters. Sume has no LTX; its route is separate clips joined on Timeline 1.0.
- LTX-2.5 quantization: fp8-cast, fp8-scaled-mm or NVFP4?
LTX-2.5's repo offers fp8-cast with bf16 checkpoints and fp8-scaled-mm on Hopper; the Hugging Face card lists NVFP4 and int8. Which flag fits which GPU.
- Lyria 3.5 is single-turn and varies per call: keep the artifact
Google says Lyria generation is single-turn, not iteratively editable and varies between calls. Why to save every Sume Music artifact you like.
Written by Sume