Run LTX-2.5 locally: Python 3.12, CUDA 12.7, three routes
The LTX-2.5 model card recommends Python 3.12+, CUDA 12.7+ and PyTorch 2.7, with a CLI, ComfyUI or Diffusers. What a hosted job on Sume replaces.

The Lightricks/LTX-2.5 model card (read 2026-10-02) recommends Python 3.12 or newer, CUDA 12.7 or newer and PyTorch about 2.7 for local runs, and offers three ways to run it: the ltx-pipelines Python CLI, ComfyUI with official workflows, or a Diffusers-compatible package. Sume does not host LTX, so a hosted equivalent means using another model from Sume's video catalog.
What the model card lists for a local run
Everything below is read from the LTX-2.5 card on 2026-10-02.
| Requirement | Listed value |
|---|---|
| Python | 3.12 or newer |
| CUDA | 12.7 or newer |
| PyTorch | About 2.7 |
| Run options | ltx-pipelines CLI, ComfyUI workflows, or Diffusers-compatible package |
| Frame count | num_frames % 8 == 1, up to 121 frames |
| Dimensions | Width and height divisible by 32 |
The hidden parts of running locally
The card gives versions, but it does not give you a service. For anything beyond a test you will also need GPU capacity sized for a 22B model, a job queue so concurrent requests do not collide, storage for outputs, retries for failed jobs, and a way to cap spend.
The frame rule matters when you plan lengths: a clip of N frames must satisfy N % 8 == 1, so valid counts go 9, 17, 25 and so on up to 121. Pick a frame rate first and work out whole-second lengths from it.
What a hosted job looks like on Sume
On Sume you do not set up GPUs. You submit POST /v1/videos with a catalog model id, receive a job id and polling URL, poll until the job is completed, then download from the content URL, as described in the video docs. Durations are whole seconds, and price is listed per model on the API pricing page.
The trade is control: you pick from the catalog's models rather than loading your own weights or LoRAs, and Sume takes no LoRA field.
When local still wins
If none of these apply, a hosted job is usually less work than maintaining a CUDA stack.
- You need custom weights or fine-tuning that a catalog does not offer.
- Footage or prompts cannot leave your network.
- Volume is high and steady enough that owned GPUs cost less than per-clip prices.
Sources
Related posts
More in Developers
- Luma Build tier: 10 concurrent Ray jobs, 20 requests a minute
Luma's Build tier allows 10 concurrent Ray video jobs, 20 requests a minute and $5000 a month. How Sume's plan concurrency and queue capacity differ.
- Luma Dream Machine API prompt rules: 3 to 5000 characters
Luma's Dream Machine API rejects prompts under 3 or over 5000 characters, and loop with keyframes. The pre-submit errors, and where Sume's checks live.
- Luma generation states and callback_url vs Sume statuses
Luma's Dream Machine API reports dreaming, completed and failed, with a callback_url POST. How that maps to Sume's queued, processing and completed.
- Luma modify video: ray-flash-2 15 seconds, ray-2 10, 100 MB
Luma's Modify Video allows 15 seconds on ray-flash-2 and 10 on ray-2, with a 100 MB source. Sume edits video with video_url on gemini-omni-flash-1.1.
Written by Sume