LTX-2.5 distilled or dev transformer: which file to download
Download the distilled bf16 transformer to generate (fixed 8-step schedule, CFG 1) and the dev one to train. int8 files are ComfyUI-only. Sume lists no LTX id.

For generating clips, download the distilled bf16 transformer from the LTX-2.5 repository; the card describes it as a fixed 8-step schedule with CFG 1. Use the dev transformer when you want to train or fine-tune, because the card calls it the full, trainable DiT. The int8 files are for ComfyUI only, and the NVFP4 file needs Blackwell hardware or ComfyUI. Sume's catalog has no LTX id, so this choice only matters if you self-host.
Everything about the files comes from the LTX-2.5 model card, read 2026-10-02. Sume's side is from Video generation.
What is in the LTX-2.5 checkpoint pack?
The card says LTX-2.5 ships as a split pack with one safetensors file per component, not one monolith, and that each CLI flag or loader points at one file. That is why the download command lists a transformer, a text encoder, two VAEs, a duration head and a spatial upscaler separately.
The text encoder is gemma4-12b-with-proj-ltx-2.5-bf16.safetensors, a Gemma 4 12B encoder plus projections. There are two video VAEs: the card calls the DiffVAE higher quality and heavier, and the Conv VAE faster and lighter.
| File | What the card says |
|---|---|
| ltx-2.5-22b-distilled-transformer-bf16 | Distilled DiT, bf16; fixed 8-step schedule, CFG 1 |
| ltx-2.5-22b-dev-transformer-bf16 | Full, trainable DiT, bf16 |
| ltx-2.5-22b-distilled-transformer-comfy-int8-convrot | Distilled, Comfy int8; ComfyUI only, not for ltx-pipelines or PyTorch |
| ltx-2.5-22b-dev-transformer-comfy-int8-convrot | Full DiT, Comfy int8; ComfyUI only |
| ltx-2.5-22b-distilled-transformer-nvfp4 | Distilled, NVFP4; ComfyUI, or ltx-pipelines with --quantization nvfp4-prequant (Blackwell, ltx-kernels) |
What do I download to run the distilled model?
The card's distilled split pack is below. It needs a Hugging Face login first, and the repository is gated, as this post on the gate explains. If you run short of VRAM, the card suggests --quantization fp8-cast and --offload cpu, and it says to use the bf16 files with ltx-pipelines because the int8 files do not load there.
hf auth login
hf download Lightricks/LTX-2.5 \
diffusion_models/ltx-2.5-22b-distilled-transformer-bf16.safetensors \
text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \
vae/ltx-2.5-video-vae-bf16.safetensors \
vae/ltx-2.5-audio-vae-bf16.safetensors \
model_patches/ltx-2.5-duration-head-bf16.safetensors \
latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors \
--local-dir models/ltx-2.5Does any of this apply on Sume?
No checkpoint choice exists on Sume. A Video Router job names a catalog id such as seedance-2.5 or wan-3.0, and the request schema reserves model_params for per-model knobs with an empty allowlist in v1. The OpenAPI model enum has no LTX row, so run LTX-2.5 yourself or call a model Sume lists. Read GET /v1/video-router/models for the current catalog.
Sources
Related posts
More in Models
- LTX-2.5 duration predictor: optional, and Sume sets seconds
LTX-2.5 has an optional duration predictor that sets the frame count from the prompt. Sume has no such mode: you send duration in whole seconds.
- Is LTX-2.5 gated on Hugging Face? Agree, log in, then download
Yes: the LTX-2.5 repository is gated. Agree to share your contact details on the model page, run hf auth login, then hf download. It is free under $10M revenue.
- MAI-Image-2.6 output cap is 2,359,296 pixels; Sume sizes differ
MAI-Image-2.6 sets a 2,359,296-pixel ceiling and a 768-pixel minimum edge. Sume sets size per model with tiers, ratios and, for GPT models, custom pixels.
- MAI-Image-2.6 edits take 5 references; Sume takes 10 or 16
MAI-Image-2.6 in Foundry accepts up to five JPEG or PNG reference images per edit. On Sume, input_references tops out at 10, or 16 on GPT Image 2.5.
Written by Sume