LTX-2.5 Ingredients LoRA: a reference sheet for consistent clips

LTX-2.5 Ingredients conditions a 121-frame clip on a character, prop and location sheet. How the sheet works and how Sume input references differ.

5 min readSume
All posts

LTX-2.5 Ingredients is an open adapter that makes a short video whose characters, props and location stay faithful to a single reference sheet you supply. On Sume you pass reference images to a hosted model through input_references instead, with no sheet layout to build, but also no promise of the adapter's fidelity.

Adapter facts are from the Ingredients card, read on 2026-10-02; Sume facts are from Video generation and Video Router.

What is a reference sheet in Ingredients?

The card calls it reference sheet conditioning. The sheet is one composite static image on a black background. It holds clean panels for each character (a face close-up and a body turnaround), for props rendered product-style, and for the location. The prompt then has two parts: one that describes the sheet and one that describes the action.

  • Control video: the sheet looped into a static video of at least 121 frames at the output resolution.
  • Output: clips at 768x448, 121 frames, 24 fps, which is about five seconds.
  • Run in ComfyUI with the specialised IC-LoRA workflow that wires in the reference input, as the card describes.

Why does the sheet layout matter?

Because the model reads the sheet as a picture, not as a list. If panels are cluttered or break the specified layout, identity can drift. The card's strict layout is the price of consistent characters across clips, and it is work for you: you have to build the sheet, loop it into a control video and write the two-part prompt each time.

How do Sume input references compare?

On POST /v1/videos, input_references carries reference images for reference-to-video, and only models whose supported_input_references lists a type accept it. Video Router adds model-specific ceilings; for Gemini Omni Flash 1.1 it is up to 10 images and 3 short videos, addressed as <IMAGE_REF_0> in the prompt. If you send both frame_images and input_references, the docs say frame images win and the request is treated as image-to-video.

There is no sheet format, and the docs make no claim that three separate photos will hold a character as tightly as a trained adapter. What you get is simplicity: send the photos, name them in the prompt, receive a job.

Reference input compared (read 2026-10-02)
ItemLTX-2.5 IngredientsSume input_references
Input shapeOne sheet, looped to 121+ framesSeparate images, optionally short videos
Output size768x448, 121 frames, 24 fpsPer model, for example 2 to 30 s on wan-3.0
SetupComfyUI workflow, two-part promptOne JSON request
Fidelity claimFaithful to the sheetReference guidance, no stated guarantee

When is the extra setup worth it?

When a series needs the same character and prop across many clips and you can afford GPU time, the sheet approach is built for that. When you need one or two clips, or you are testing a concept, a hosted reference call is faster. The output size is also worth a look: 768x448 is small, so plan an upscale step, and see the Sume reference limits post for what hosted models accept.

Sources

Related posts

More in Models

All Models posts

Written by Sume