LTX-2.5 quantization: fp8-cast, fp8-scaled-mm or NVFP4?

LTX-2.5's repo offers fp8-cast with bf16 checkpoints and fp8-scaled-mm on Hopper; the Hugging Face card lists NVFP4 and int8. Which flag fits which GPU.

5 min readSume
All posts

For LTX-2.5 the GitHub repository documents two FP8 modes: fp8-cast for bf16 checkpoints when memory is tight, and fp8-scaled-mm for Hopper-class GPUs with native FP8 and fp8 checkpoints. The Hugging Face card also lists NVFP4 and ComfyUI int8 quantized files. I could not confirm a single VRAM figure for each mode, so pick by GPU generation, then measure.

What does each mode do?

From the LTX-2 README: FP8-cast lowers the memory footprint and is used with bf16 checkpoints. FP8-scaled-mm runs on Hopper and newer GPUs with native FP8 support and is used with fp8 checkpoints. Plain bf16 runs without quantization. When GPU memory is the constraint, the README shows --quantization fp8-cast together with --offload cpu or --offload disk.

LTX-2.5 quantization options (read 2026-10-02)
OptionWhere documentedPairs with
bf16, no quantizationLTX-2 READMEbf16 checkpoint
fp8-castLTX-2 READMEbf16 checkpoint, optional CPU or disk offload
fp8-scaled-mmLTX-2 READMEfp8 checkpoint on Hopper or newer
NVFP4LTX-2.5 Hugging Face cardListed as an option
int8 (ComfyUI)LTX-2.5 Hugging Face cardComfyUI-specific quantized files

How big is the download?

The README says the model package is about 66 GB, split into component files: transformer, text encoder, VAEs and upscalers. The 22B distilled transformer is the fast path and the dev transformer is the trainable one; see distilled or dev for that choice. Resolution must be divisible by 32 and the frame count must satisfy num_frames % 8 == 1.

Which should I try first?

Start from your card. On a Hopper-generation GPU, try fp8 checkpoints with fp8-scaled-mm. On an older or smaller card, try fp8-cast with offload and accept slower runs. Use ComfyUI int8 or NVFP4 files only if you work in that stack and your GPU supports them; the card does not say which GPUs, so check the files' notes.

Whatever you pick, render the same prompt at bf16 once if you can, and compare. Quantization changes output, and I have no benchmark to tell you by how much.

When should I skip local quantization entirely?

If you are tuning flags to squeeze a model onto a card for a client deliverable, a hosted route may be cheaper than your time. Sume's docs do not list an LTX id in the pages I read, so confirm with GET /v1/videos/models and see what Sume lists instead of an LTX id. The Video generation docs explain the job flow for whichever id you do find.

How do I compare quantized output fairly?

Fix everything except the quantization mode. Use the same prompt, the same seed (the example command in the README passes --seed 42), the same frame count (121 in the example) and the same checkpoint family. Render each mode, then compare motion, text rendering and audio sync, not just a single frame.

Keep notes on peak memory and wall-clock time for each run. Those two numbers, from your own card, are the only ones that matter for your decision, and they are the ones this post cannot give you.

Sources

Related posts

More in Models

All Models posts

Written by Sume