LTX-2.5 on Windows or Mac: natten, attention and fallbacks
LTX-2's README: natten (VAE decode) is Linux and CUDA only, with a Triton or eager fallback elsewhere; FlashAttention 4 on B200, 3 on Hopper, SDPA otherwise.

LTX-2.5 can be installed on Windows or macOS, but its fastest paths are for Linux with CUDA. The LTX-2 README says the natten extra, the fastest backend for the diffusion video VAE, is Linux and CUDA only, and that on Windows and macOS it is skipped automatically and decoding falls back to a Triton or eager implementation. The model card still lists Python 3.12 or higher and CUDA 12.7 or higher, so a Mac without CUDA is outside the card's stated requirements.
Which backend does each part get?
| Part | Hardware | README recommendation |
|---|---|---|
| Attention | B200 (datacenter Blackwell) | FlashAttention 4, installed manually |
| Attention | Hopper | FlashAttention 3 wheel |
| Attention | Other CUDA GPUs | PyTorch SDPA, used automatically |
| Video VAE decode | Linux with CUDA | optional natten extra |
| Video VAE decode | Windows or macOS | Triton or eager fallback |
What does a fallback cost?
The README does not give speeds for the fallbacks, so I cannot say how much slower they are. Expect to measure your own. An installed attention backend is selected automatically; forcing one is a Python-API option, not a CLI flag. The memory flags, --quantization fp8-cast and --offload cpu or disk, trade speed for memory.
Is WSL the way out?
The README does not say, so this is a suggestion rather than a vendor claim: Windows users with an NVIDIA GPU often run Linux tooling inside WSL, which gives the Linux and CUDA path the natten extra expects. Check that the CUDA version inside it meets the card's 12.7 floor. On a Mac there is no CUDA, so you are outside the card's stated requirements either way.
Before you spend a day on setup, run the smallest possible job end to end, and time it. If the fallback is too slow for your use, a hosted job is the shorter path.
When is the hosted route the simpler call?
If the machine is a Mac or a Windows laptop and the goal is a clip rather than a pipeline, send the job to a hosted id and download the MP4. Sume does not list LTX (catalog code, read 2026-10-09), so pick from GET /v1/videos/models. A request is plain JSON with model, prompt, duration, resolution, so it runs the same from any OS. See the video docs.
Sources
Related posts
More in Developers
- LTX-2.5 pipelines: Distilled, DFR or two-stage, which to run
The LTX-2 repo names eleven pipelines for LTX-2.5. Which one is fastest, which is guided, which does keyframes or audio, and where the Sume fields line up.
- Lyria 3.5 has no duration field: how to ask for a 45-second track
Sume's music router rejects duration and duration_seconds. Write the length into the prompt, with section timestamps, and pay a flat $0.125 per generation.
- MCP authorization spec: 6 requirements vs what Sume documents
The MCP authorization spec asks for resource metadata, PKCE and a resource parameter. Here is what hosted Sume MCP documents for each, and what stays open.
- Sume hosted MCP from a CI runner: API key or OAuth?
A headless runner cannot finish the OAuth consent page. Hosted Sume MCP also takes an API key as Bearer or x-api-key. What that session can call, and the rules.
Written by Sume