LTX-2.5 Layout-To-Render: Blender playblast to a finished shot
LTX-2.5 Layout-To-Render keeps a 3D playblast's camera path and layout and restyles it from a reference image. What Sume's video references can and cannot do.

LTX-2.5 Layout-To-Render is an open adapter that turns a grey clay render or blocky playblast from Blender or Unreal into a finished video shot, keeping the camera movement and object placement and taking its look from a reference image. Sume's hosted video models accept reference videos, but its docs do not promise that a reference's camera path or layout is preserved, so this is not the same job.
The adapter details come from the Layout-To-Render card, read on 2026-10-02. Sume facts come from Video generation and Video Router.
What goes into Layout-To-Render?
Three inputs, per the card. First, a layout video: a 3D viewport animation or playblast, preferably at 24 fps with dimensions divisible by 64. Second, one or more reference images that carry the art direction, derived from the first frame of the clay render. Third, a one- or two-sentence text prompt describing the finished shot.
- Output: a finished video with the layout's camera path and object positions, restyled in lighting, materials and palette.
- Width and height divisible by 64; the card says 1920x1088 is a known good size.
- Frame count snaps to 8x k plus 1.
- Runs through ComfyUI with the ComfyUI-LTXVideo nodes at LoRA strength 1.0, on the distilled transformer, text encoder and VAEs from LTX-2.5.
Why is a layout-conditioned adapter useful?
A previs or animation team already knows the camera move and the blocking. The expensive part is the final look. A layout-conditioned adapter keeps the blocking fixed while a style frame sets the look, so you iterate on style without re-animating. A prompt-only text-to-video model cannot hold a camera path that exact, because a prompt describes a move in words.
What does Sume offer with a reference video?
The docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, and that Gemini Omni Flash 1.1 takes video references but not audio. In Video Router, Gemini Omni Flash 1.1 takes up to 10 reference images and up to 3 reference videos of at most 3 seconds each, addressed in the prompt as <IMAGE_REF_0> and <VIDEO_REF_0>. Its edit mode takes one video_url and a prompt for the change, 3 to 10 seconds, with no layout-to-render concept.
A reference in those routes guides the output. The docs do not describe it as locking camera path or object placement frame by frame, so do not assume it does. A blocky clay playblast sent as a reference may influence the result, but test it on your own shot before promising a client the move will match.
| Need | LTX-2.5 Layout-To-Render | Sume video references |
|---|---|---|
| Keep the 3D camera path | Designed to | Not documented as a guarantee |
| Style from an image | Reference image | Image references on supporting models |
| Where it runs | ComfyUI on your GPU | Hosted job |
| Clip limit | Frame count 8k+1, native 1920x1088 | Model limits, for example Gemini Omni 3 to 10 s |
Which route fits a previs pipeline?
If the camera move is fixed and you own the GPU, Layout-To-Render is the only one here built for it. If you are exploring looks from a script, hosted text and reference generation is quicker to start and needs no setup. A realistic split is to explore the look on a hosted model, then lock the shot with the adapter. Sume lists no LTX id; check GET /v1/videos/models before you write any code that assumes otherwise. The licence on the adapter is the LTX-2 community licence, so read the terms for your company.
Sources
Related posts
More in Models
- LTX-2.5 native multishot: one prompt, or clips joined on Sume
LTX-2.5 adds native multishot: connected scenes in one pass, consistent characters. Sume has no LTX; its route is separate clips joined on Timeline 1.0.
- LTX-2.5 Pixel Spatial Upscaler: draft at 280p, then upscale 2x
The LTX-2.5 Pixel Spatial Upscaler re-renders a low-res draft 2x with invented detail. How the draft workflow works and how Sume's draft-then-final differs.
- LTX-2.5 quantization: fp8-cast, fp8-scaled-mm or NVFP4?
LTX-2.5's repo offers fp8-cast with bf16 checkpoints and fp8-scaled-mm on Hopper; the Hugging Face card lists NVFP4 and int8. Which flag fits which GPU.
- LTX-2.5 Refine Details: 22 minutes per clip on an RTX PRO 6000
The LTX-2.5 Refine Details LoRA rebuilds texture up to 4K but takes about 22 minutes for 121 frames on an RTX PRO 6000. A hosted upscale compared.
Written by Sume