Run LTX-2.5 locally: Python 3.12, CUDA 12.7, three routes

The LTX-2.5 model card recommends Python 3.12+, CUDA 12.7+ and PyTorch 2.7, with a CLI, ComfyUI or Diffusers. What a hosted job on Sume replaces.

5 min readSume
All posts

The Lightricks/LTX-2.5 model card (read 2026-10-02) recommends Python 3.12 or newer, CUDA 12.7 or newer and PyTorch about 2.7 for local runs, and offers three ways to run it: the ltx-pipelines Python CLI, ComfyUI with official workflows, or a Diffusers-compatible package. Sume does not host LTX, so a hosted equivalent means using another model from Sume's video catalog.

What the model card lists for a local run

Everything below is read from the LTX-2.5 card on 2026-10-02.

LTX-2.5 local setup, vendor facts (read 2026-10-02)
RequirementListed value
Python3.12 or newer
CUDA12.7 or newer
PyTorchAbout 2.7
Run optionsltx-pipelines CLI, ComfyUI workflows, or Diffusers-compatible package
Frame countnum_frames % 8 == 1, up to 121 frames
DimensionsWidth and height divisible by 32

The hidden parts of running locally

The card gives versions, but it does not give you a service. For anything beyond a test you will also need GPU capacity sized for a 22B model, a job queue so concurrent requests do not collide, storage for outputs, retries for failed jobs, and a way to cap spend.

The frame rule matters when you plan lengths: a clip of N frames must satisfy N % 8 == 1, so valid counts go 9, 17, 25 and so on up to 121. Pick a frame rate first and work out whole-second lengths from it.

What a hosted job looks like on Sume

On Sume you do not set up GPUs. You submit POST /v1/videos with a catalog model id, receive a job id and polling URL, poll until the job is completed, then download from the content URL, as described in the video docs. Durations are whole seconds, and price is listed per model on the API pricing page.

The trade is control: you pick from the catalog's models rather than loading your own weights or LoRAs, and Sume takes no LoRA field.

When local still wins

If none of these apply, a hosted job is usually less work than maintaining a CUDA stack.

  • You need custom weights or fine-tuning that a catalog does not offer.
  • Footage or prompts cannot leave your network.
  • Volume is high and steady enough that owned GPUs cost less than per-clip prices.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume