LTX-2.5 pipelines: Distilled, DFR or two-stage, which to run
The LTX-2 repo names eleven pipelines for LTX-2.5. Which one is fastest, which is guided, which does keyframes or audio, and where the Sume fields line up.

For a first LTX-2.5 run, use DistilledPipeline: the repository describes it as the fastest text/image-to-video option, started with python -m ltx_pipelines.distilled. Choose DFRPipeline for production-quality generation with spatial detailing, and TI2VidTwoStagesPipeline when you want guided generation with CFG and STG.
What pipelines does the README list?
| Pipeline | What the README says it is for |
|---|---|
| DistilledPipeline | fastest text/image-to-video |
| DFRPipeline | production-quality generation with spatial detailing |
| TI2VidTwoStagesPipeline | guided two-stage, CFG and STG |
| ICLoraPipeline | video-to-video transformations |
| KeyframeInterpolationPipeline | keyframe interpolation |
| A2VidPipelineTwoStage | audio-to-video |
| DubItPipeline | speaker identity matching with lip sync |
| RetakePipeline | region-specific video regeneration |
| TI2VidTwoStagesHQPipeline | same guided two-stage flow with the res_2s sampler (fewer steps) |
| TI2VidOneStagePipeline | single-stage generation for quick prototyping |
| HDRICLoraPipeline | video-to-video SDR to HDR |
What else does the README set?
The README lists an optional duration head, ltx-2.5-duration-head-bf16.safetensors, so you can omit --num-frames and have length predicted from the prompt. Memory flags are --quantization fp8-cast and --offload to cpu or disk. It says gradient estimation reduces steps from 40 to 20 to 30. A prompt enhancer is on by an enhance_prompt parameter; the README gives no further detail on it, so test with it on and off.
What are the nearest hosted fields?
Sume does not list LTX (catalog code, read 2026-10-09), so these are the nearest request shapes on listed models, not equivalents.
- Keyframes:
frame_imageswithfirst_frameandlast_frame, on models whosesupported_frame_imagesinclude both. - Audio-driven:
input_referenceswith anaudio_url, on models that list audio references such asminimax-h3andwan-3.0. - Edit an existing clip: the Video Router
video_urledit ongemini-omni-flash-1.1(whole-clip edit, not a region retake). - Guided steps, CFG or STG: not exposed.
What is a sensible order to try them?
Run DistilledPipeline first to prove the install and get a baseline time. Move to DFRPipeline or TI2VidTwoStagesPipeline for the clips that matter, and compare against the distilled output on the same prompt. Use the specialised pipelines (keyframes, audio-to-video, dubbing, retake) only when the task calls for them; each one has its own inputs. For example, A2VidPipelineTwoStage needs an audio file, and KeyframeInterpolationPipeline needs the keyframes you want to hit, so gather those before you start.
Write the pipeline name, the checkpoint and the flags beside every output. With eleven pipelines and several checkpoints, that note is what keeps your tests comparable.
Check each id's fields at GET /v1/videos/models before you build; the video docs list them.
Sources
Related posts
More in Developers
- Lyria 3.5 has no duration field: how to ask for a 45-second track
Sume's music router rejects duration and duration_seconds. Write the length into the prompt, with section timestamps, and pay a flat $0.125 per generation.
- MCP authorization spec: 6 requirements vs what Sume documents
The MCP authorization spec asks for resource metadata, PKCE and a resource parameter. Here is what hosted Sume MCP documents for each, and what stays open.
- Sume hosted MCP from a CI runner: API key or OAuth?
A headless runner cannot finish the OAuth consent page. Hosted Sume MCP also takes an API key as Bearer or x-api-key. What that session can call, and the rules.
- MCP generate_video max_spend_usd for Wan 3.0: 480p, 720p, 1080p
What to put in max_spend_usd when an MCP client calls generate_video for Wan 3.0: $0.625 to $7.50 for 10 and 30 seconds. Includes the dry_run body.
Written by Sume