LTX-2.5 native multishot: one prompt, or clips joined on Sume
LTX-2.5 adds native multishot: connected scenes in one pass, consistent characters. Sume has no LTX; its route is separate clips joined on Timeline 1.0.

LTX-2.5's model card on Hugging Face (read 2026-10-02) says the release introduces native multishot generation: connected scenes in a single pass, keeping character identity, environment and visual style across cuts. Sume does not list an LTX model, so the Sume route to a multi-shot video is to generate each shot as its own clip and assemble them with Timeline 1.0.
What the model card says
The Lightricks/LTX-2.5 card (read 2026-10-02) lists the items below. Other coverage dates the launch to Aug 11, 2026; this page relies on the card itself for the specs.
| Item | Listed on the model card |
|---|---|
| Transformer | 22B distilled DiT |
| Text encoder | Gemma 4 12B |
| Multishot | Native, connected scenes in a single pass |
| Frame count | num_frames % 8 == 1, up to 121 frames at standard frame rates |
| Resolution | Divisible by 32; from 544x960 (stage 1) up to 4K with spatial upscaling |
| Commercial use | Free under $10M annual revenue under the LTX-2.x Community License; above that a paid agreement |
One pass versus one shot at a time
A single multishot generation has one advantage: the model sees the whole sequence, so a character who appears in shots one and three is more likely to match. The cost is control. If shot two is wrong you regenerate the sequence, or use whatever per-segment tools the pipeline offers.
Generating shot by shot reverses that trade. You can fix one clip without touching the others, but you carry consistency yourself, with references, a fixed first frame, or a shared prompt block.
The Sume route
Sume's video API generates a clip per job. Limits vary by model: for example Wan 3.0 accepts 2 to 30 seconds and Seedance 2.5 4 to 30 seconds per the video docs. To join shots, Timeline 1.0 takes one audio spine plus ordered video[] slots and returns one MP4, handling sequencing and transitions.
Models that accept image references, such as those listed in the catalog, can help keep a character stable across separate shots. Sume has no single-pass multishot switch, and does not claim cross-shot consistency from the join itself.
Decision notes
Check the model card again before shipping; specs and license terms can change.
- Short sequence, one look, you control the GPU: LTX-2.5 multishot may suit.
- Many iterations on single shots, hosted: separate clips plus a timeline.
- Revenue above $10M: read the LTX license terms before you build on the weights.
Sources
Related posts
More in Models
- Lyria 3.5 is single-turn and varies per call: keep the artifact
Google says Lyria generation is single-turn, not iteratively editable and varies between calls. Why to save every Sume Music artifact you like.
- Lyria 3.5 in another language: prompt in it, then check the song
Google says Lyria 3.5 makes music in other languages when you prompt in that language. On Sume, write the brief and lyrics in it, then check the result.
- MAI-Image-2.6 output cap is 2,359,296 pixels; Sume sizes differ
MAI-Image-2.6 sets a 2,359,296-pixel ceiling and a 768-pixel minimum edge. Sume sets size per model with tiers, ratios and, for GPT models, custom pixels.
- MAI-Image-2.6 edits take 5 references; Sume takes 10 or 16
MAI-Image-2.6 in Foundry accepts up to five JPEG or PNG reference images per edit. On Sume, input_references tops out at 10, or 16 on GPT Image 2.5.
Written by Sume