LTX-2.5 quantization: fp8-cast, fp8-scaled-mm or NVFP4?
LTX-2.5's repo offers fp8-cast with bf16 checkpoints and fp8-scaled-mm on Hopper; the Hugging Face card lists NVFP4 and int8. Which flag fits which GPU.

For LTX-2.5 the GitHub repository documents two FP8 modes: fp8-cast for bf16 checkpoints when memory is tight, and fp8-scaled-mm for Hopper-class GPUs with native FP8 and fp8 checkpoints. The Hugging Face card also lists NVFP4 and ComfyUI int8 quantized files. I could not confirm a single VRAM figure for each mode, so pick by GPU generation, then measure.
What does each mode do?
From the LTX-2 README: FP8-cast lowers the memory footprint and is used with bf16 checkpoints. FP8-scaled-mm runs on Hopper and newer GPUs with native FP8 support and is used with fp8 checkpoints. Plain bf16 runs without quantization. When GPU memory is the constraint, the README shows --quantization fp8-cast together with --offload cpu or --offload disk.
| Option | Where documented | Pairs with |
|---|---|---|
| bf16, no quantization | LTX-2 README | bf16 checkpoint |
| fp8-cast | LTX-2 README | bf16 checkpoint, optional CPU or disk offload |
| fp8-scaled-mm | LTX-2 README | fp8 checkpoint on Hopper or newer |
| NVFP4 | LTX-2.5 Hugging Face card | Listed as an option |
| int8 (ComfyUI) | LTX-2.5 Hugging Face card | ComfyUI-specific quantized files |
How big is the download?
The README says the model package is about 66 GB, split into component files: transformer, text encoder, VAEs and upscalers. The 22B distilled transformer is the fast path and the dev transformer is the trainable one; see distilled or dev for that choice. Resolution must be divisible by 32 and the frame count must satisfy num_frames % 8 == 1.
Which should I try first?
Start from your card. On a Hopper-generation GPU, try fp8 checkpoints with fp8-scaled-mm. On an older or smaller card, try fp8-cast with offload and accept slower runs. Use ComfyUI int8 or NVFP4 files only if you work in that stack and your GPU supports them; the card does not say which GPUs, so check the files' notes.
Whatever you pick, render the same prompt at bf16 once if you can, and compare. Quantization changes output, and I have no benchmark to tell you by how much.
When should I skip local quantization entirely?
If you are tuning flags to squeeze a model onto a card for a client deliverable, a hosted route may be cheaper than your time. Sume's docs do not list an LTX id in the pages I read, so confirm with GET /v1/videos/models and see what Sume lists instead of an LTX id. The Video generation docs explain the job flow for whichever id you do find.
How do I compare quantized output fairly?
Fix everything except the quantization mode. Use the same prompt, the same seed (the example command in the README passes --seed 42), the same frame count (121 in the example) and the same checkpoint family. Render each mode, then compare motion, text rendering and audio sync, not just a single frame.
Keep notes on peak memory and wall-clock time for each run. Those two numbers, from your own card, are the only ones that matter for your decision, and they are the ones this post cannot give you.
Sources
Related posts
More in Models
- Lyria 3.5 is single-turn and varies per call: keep the artifact
Google says Lyria generation is single-turn, not iteratively editable and varies between calls. Why to save every Sume Music artifact you like.
- Lyria 3.5 in another language: prompt in it, then check the song
Google says Lyria 3.5 makes music in other languages when you prompt in that language. On Sume, write the brief and lyrics in it, then check the result.
- MAI-Image-2.6 output cap is 2,359,296 pixels; Sume sizes differ
MAI-Image-2.6 sets a 2,359,296-pixel ceiling and a 768-pixel minimum edge. Sume sets size per model with tiers, ratios and, for GPT models, custom pixels.
- MAI-Image-2.6 edits take 5 references; Sume takes 10 or 16
MAI-Image-2.6 in Foundry accepts up to five JPEG or PNG reference images per edit. On Sume, input_references tops out at 10, or 16 on GPT Image 2.5.
Written by Sume