LTX-2 diffusion decoder or convolutional decoder: which to use

LTX-2 ships a diffusion decoder (better quality, more VRAM) and a lighter convolutional one. What the README says, a draft-then-final habit, and hosted jobs.

5 min readSume
All posts

LTX-2's repository ships two video decoders: a diffusion decoder that gives better quality but needs more VRAM and decode time, and a lighter convolutional decoder with no extra dependencies. Pick the diffusion decoder for final shots and the convolutional one for drafts, and expect the choice to matter most at the end of the pipeline, not in sampling.

Details come from the LTX-2 README, the repository and the LTX-2.5 card, all read 2026-10-03. Our other LTX-2.5 posts cover checkpoints, quantization and upscalers; this one is about the decoder.

What are the two decoders?

The README lists two video VAE options. The standard one is a diffusion decoder, named NADiffusionDecoder, described as improved quality at the cost of longer decode time and more VRAM. It is fastest with the optional natten extra and falls back to Triton or eager neighborhood attention without it. The alternative is a convolutional decoder, described as lighter and needing no extra dependencies. The LTX-2.5 card calls the diffusion decoder new and says it replaces VAE reconstruction.

The README gives no decode time or memory figure for either option, so the only numbers you will get are your own.

A practical way to record your own numbers: for one clip, log the decode time and peak VRAM with each decoder, with the same latents and the same card, and keep the pair in your project notes. Those two figures, not the README's adjectives, tell you whether the diffusion decoder fits your machine for a batch of finals.

LTX-2 decoder options from the README, read 2026-10-03.
DecoderQualityCostDependencies
NADiffusionDecoder (diffusion)Improved qualityLonger decode time, more VRAMFastest with the natten extra; Triton or eager fallback
ConvolutionalBaseline, lighterLighterNone extra

What else does the repository add for final quality?

The README also lists a DFRPipeline for production-quality text and image to video, described as slower and heavier on VRAM: the same distilled transformer, extra generated keyframes, a spatial detailing pass and optional 2x or 4x frame rate. The LTX-2.5 card separately lists diffusion fidelity rendering that allocates compute by scene complexity. The README does not say the two are the same thing, so treat them as related claims until you test.

The pattern in all of them is the same: spend the extra compute at the end, on keyframes, detail and decode, and keep the first pass cheap. Our post on the pixel spatial upscaler draft workflow covers the draft-first side.

How do you use the choice in a workflow?

Use a two-pass habit. Render drafts with the convolutional decoder, which is lighter and needs no extra dependencies, and keep the seed and prompt. When a draft is approved, re-run with the diffusion decoder, and install the natten extra first, since the README says it is fastest with it. Compare the same frame from both on a large monitor before you decide the extra time is worth it; the README states the direction of the trade, not its size.

On a card with little VRAM, decode is where an out-of-memory error can appear after a long sampling run, because the diffusion decoder uses more memory. If that happens, keep the convolutional decoder for that clip and fix the quality elsewhere.

Where does a hosted job fit?

A hosted route hides the decoder choice. Sume's Video generation docs list only the models in the live catalog, with their resolutions, durations and reference types, and GET /v1/videos/models is the place to check whether any open model you care about is routable; send only ids it lists. The request has no decoder field.

For a clip you already have, Sume has a separate paid video_upscale_create tool in the hosted MCP inventory; the pages I read do not describe how it renders, so do not assume it works like a diffusion decoder. Judge it on your own clip.

Use local LTX-2 when you want control of the decode and own the GPU. Use a hosted job when you want a finished clip without managing dependencies, and accept that the decoder is not your choice.

Sources

Related posts

More in Models

All Models posts

Written by Sume