LTX keyframe interpolation is open-source only; Sume takes last_frame
LTX's KeyframeInterpolationPipeline has no API endpoint. On Sume, frame_images accepts first_frame and last_frame on models that report support in the catalog.

LTX's keyframe interpolation is open-source only: the post says KeyframeInterpolationPipeline has no API endpoint. If you want a clip that starts on one known image and ends on another through a hosted API, Sume's frame_images takes a first_frame and a last_frame entry on models whose catalog row reports support.
LTX's statement is from its post How to make longer AI videos, read 2026-10-01. Sume's side is from Video generation and the Video Router docs.
What does LTX say about keyframe interpolation?
The post presents it as the method for getting from a known first image to a known last image, interpolating each gap as its own segment, and each segment must still satisfy the 8k+1 frame rule. It states that this one is open-source only. The hosted LTX API offers extend and retake endpoints for continuation, not this.
How do I send a first and last frame on Sume?
Use frame_images: each entry has a frame_type of first_frame or last_frame. Entries trigger image-to-video; input_references is the separate reference-to-video route, and if both are sent frame_images takes precedence. Check the catalog field supported_frame_images first, which reports which frame_type values a model accepts.
Which catalog models list an end frame?
From the Video Router table read 2026-10-01:
| Model id | Listed capability | Duration |
|---|---|---|
wan-3.0 | t2v / i2v+end / r2v | 2-30s |
minimax-h3 | t2v / i2v+end / r2v | 5-15s |
minimax-h3-max | t2v / i2v+end / r2v | 5-15s |
gemini-omni-flash-1.1 | t2v / i2v+end / r2v / edit | 3-10s |
Is that the same as LTX interpolation?
Not necessarily. The docs call it image-to-video with a start and end image, and do not claim the exact segment-by-segment interpolation LTX describes. Test one pair of frames before building a chain, and use the model's own duration limits. More on one model: MiniMax H3 first and last frame. For which field wins when you mix inputs: frame_images and input_references together.
Sources
Related posts
More in Developers
- LTX retake vs Sume: fix one section of a video
LTX's POST /v1/retake regenerates one time region of a video. Sume's video edit takes a whole clip, so trim the bad section, edit it, and rejoin it.
- LTX frame counts must be 8k+1; Sume durations are whole seconds
LTX frame counts must satisfy (F-1) % 8 == 0, so 30 frames is invalid. On Sume you send whole seconds and read the allowed set from the catalog, not frames.
- Luma API 429 requests per minute: sliding window vs Sume
Luma counts requests in a sliding 60-second window and returns 429 if RPM or concurrent jobs fails. Sume returns 429 rate_limited: back off, reuse the key.
- Luma API concurrent jobs limit vs Sume plan concurrency
Luma caps active generations per API client and answers 429 when full. Sume ties concurrency to your plan and queues extra jobs until queue_full.
Written by Sume