Which AI video models take an end frame or reference clips on Sume
Sume's catalog flags show who accepts text-to-video, an end frame, reference images, clips and audio, and who edits video. Grok is image-to-video only.

Seedance 2.5, Seedance 2.0, Wan 3.0, MiniMax H3 and H3 Max take text-to-video, image-to-video, an end frame and reference images, clips and audio, while Kling Video v3 Pro takes no references and Grok Imagine Video 1.5 is image-to-video only. Among these, Gemini Omni Flash 1.1 is the one that edits an existing video, and it takes no reference audio. Special-purpose models such as h3-max-recast and higgsfield-genjutsu also take a source video.
Capability flags from the catalog
| Model | Text to video | Image to video | End frame | Reference images and videos | Reference audio | Edit a video |
|---|---|---|---|---|---|---|
| Seedance 2.5, Seedance 2.0 | yes | yes | yes | yes | yes | no |
| Wan 3.0 | yes | yes | yes | yes | yes | no |
| MiniMax H3, H3 Max | yes | yes | yes | yes | yes | no |
| Kling Video v3 Pro | yes | yes | yes | no | no | no |
| Gemini Omni Flash 1.1 | yes | yes | yes | yes | no | yes |
| Grok Imagine Video 1.5 | no | yes | no | no | no | no |
Limits on references
Each limit is quoted from the model's constraints in the catalog, and each applies per job.
- H3: up to 9 reference images, 3 clips of 2 to 15 seconds and 3 audio files, 12 files in all.
- Omni: up to 10 reference images and 3 reference clips of 3 seconds each, plus an edit mode that takes a prompt and a video_url.
- Wan 3.0: up to 10 reference images, 5 clips (15 seconds total) and 5 audio files (15 seconds total).
How to use the matrix
Start from the feature, not the brand. If you need to hold a character across shots with reference images, rule out Kling and Grok first. If you need to change an existing clip, Omni is the general choice among these, and the special-purpose edit models are the other route. If you need the last frame to land on a given image, use any model with the end-frame flag. Then pick among the survivors on price and duration limits.
A caution on limits
The flags say a feature exists, not that every combination is allowed. H3, for example, caps the total at 12 reference files and does not allow audio as the only reference, and Omni treats one reference image with no first or end frame as reference-to-video. Read the constraints list for the model before you build a form around it. Duration and resolution limits differ by model as well, so a feature that exists on two models may fit your clip on only one of them.
For migration work, the safest approach is to list the features your current clips use, tick them against this table, and test one clip per surviving model before you commit.
Sources
Related posts
More in Models
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
- Music generation API: the Sume Music Router with Lyria 3.5
Sume's Music Router turns a text prompt into a track via POST /v1/music-router/generate. sume/music-auto picks the engine, Lyria 3.5 today.
Written by Sume