LTX-2.5 quick start downloads 66 GiB: the five files and a disk plan

The LTX-2 README quick start pulls about 66 GiB for five LTX-2.5 files, including a 12B Gemma text encoder. What to budget, and what Sume lists instead.

5 min readSume
All posts

The LTX-2 README quick start downloads roughly 66 GiB for LTX-2.5: five files, led by a 22B distilled transformer in bf16 and a Gemma 4 12B text encoder. 66 GiB is about 70.9 GB in decimal units (66 x 1.0737). Plan for more than that on disk, because the README also lists a pixel-space diffusion VAE and other checkpoints for other pipelines. Sume does not list LTX-2.5 in its video catalog, so for Sume users the local route means a separate stack. Read 2026-10-08.

The five files

The hf download command in the quick start names five items under Lightricks/LTX-2.5 and writes them in folder layout under models/ltx-2.5.

LTX-2.5 quick-start downloads, README read 2026-10-08
FolderFileRole
diffusion_modelsltx-2.5-22b-distilled-transformer-bf16.safetensorsDistilled transformer, 22B
text_encodersgemma4-12b-with-proj-ltx-2.5-bf16.safetensorsText encoder with projection
vaeltx-2.5-video-vae-bf16.safetensorsVideo decoder
vaeltx-2.5-audio-vae-bf16.safetensorsAudio decoder
latent_upscale_modelsltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors2x spatial upscaler

Other things the README mentions

The README says that if you hit a 401 or 403, accept the model terms on Hugging Face and log in with a read token; fine-grained tokens need the 'read gated repos' scope. For tight memory it suggests --quantization fp8-cast --offload cpu or disk. It also describes a slower DFR path for production quality that needs more VRAM, and it notes the fastest VAE backend (natten) is Linux plus CUDA only.

The license changed on August 11, 2026. Entities with annual revenue of at least $10,000,000 need a paid license for non-exempt use, so a download is not the end of the legal work.

Disk and network plan

Budget for the weights and for headroom:

  • 66 GiB for the quick-start set, about 70.9 GB decimal.
  • Extra room for any second checkpoint, such as the dev transformer, if you switch pipelines.
  • A read token for Hugging Face and an accepted gate before the first download.
  • A Linux CUDA machine if you want the fastest decoder.

What Sume offers instead

Sume's video catalog lists ids such as wan-3.0, minimax-h3 and seedance-2.5; it does not list LTX. A hosted Wan 3.0 clip at 720p costs $0.125 per second, so ten seconds is $1.25. That removes the download, the token, and the GPU, and replaces them with a per-second price. It does not replace LTX's own audio-video model or its LoRAs; Sume's request table has no LoRA field.

Checks after the download

Verify file sizes against the Hugging Face listing, because a gated repository can return an HTML error page that looks like a model file. Check the license date shown in the repository you downloaded from: the August 11, 2026 license applies to LTX-2.5 versions released since that date. Keep the folder layout from the README (diffusion_models, text_encoders, vae, latent_upscale_models) so the loader finds each file.

A 22B transformer in bf16 is large by itself; the README points to fp8-cast and offload options for tight memory, which trade speed for fit.

Sources

Related posts

More in Models

All Models posts

Written by Sume