Which AI video models take an end frame in Sume's video picker?
Wan 3.0, MiniMax H3, H3 Max and Auto take an end frame in Sume's Videos panel; Kling 3.0 and Grok Imagine do not. How to check the API list too.

In Sume's Agents Videos panel, Auto, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept an end frame; Kling 3.0 and Grok Imagine do not, so pick one of the first four when your clip must land on a specific last image. The panel hides the end-frame tile for models that cannot use it, which is the quickest way to see the rule.
This post covers the panel's model list as it ships on main, plus the API-side check, because the two can differ. Sume's video generation docs describe the API; the panel's per-model limits live in the app code.
Which picker models accept a last frame?
The panel lists seven entries: Auto, the retiring Sume Video 1.0 alias, Kling 3.0, Wan 3.0, MiniMax H3, MiniMax H3 Max and Grok Imagine. Each carries a flag for end-frame support. Video 1.0 behaves as Auto because it is a compatibility alias for the same pipe.
| Model | Vendor | End frame slot | Reference slots |
|---|---|---|---|
| Auto | Sume | Yes | Yes |
| Kling 3.0 | Kuaishou | No | Yes |
| Wan 3.0 | Alibaba | Yes | Yes |
| MiniMax H3 | MiniMax | Yes | Yes |
| MiniMax H3 Max | MiniMax | Yes | No |
| Grok Imagine | xAI | No | No |
Why is the end frame missing for Kling 3.0?
The panel's Kling 3.0 entry is built without an end-frame slot, and the code comment says the product path for Kling create does not expose a tail frame. That is a statement about Sume's panel, not about Kling: Kling's own 4.0 versus 3.0 comparison says Kling 3.0 does not support multiple keyframes and Kling 4.0 supports up to 10. Sume lists Kling 3.0 today, so a multi-keyframe workflow on Kling is not something you can run on Sume.
Grok Imagine is image-to-video from a start frame only. xAI's video generation guide says a still image is the starting point, and the panel does not send an end image for it.
What does the API say?
The panel is a front end. For the API, ask the catalog what each model accepts instead of guessing. Every entry in GET /v1/videos/models carries supported_frame_images, which lists first_frame and/or last_frame, and a request with a model that does not list last_frame should not send one. Note that the Video Router catalog in the API lists kling-3 with end-frame support even though the panel hides the tile, so the panel is the stricter of the two.
Frame images go in the frame_images array with a frame_type. If you send both frame_images and input_references, the docs say frame_images takes precedence and the request is treated as image-to-video.
curl "https://api.sume.com/v1/videos/models" \
-H "Authorization: Bearer $SUME_API_KEY"
# read supported_frame_images on the model you plan to pinHow do I choose when I need both frames?
Start from the shot, not the model. For a product reveal that must end on a packshot, pick Wan 3.0 or MiniMax H3 and upload both frames. For a seamless loop, use the same image as first and last frame. If you also need references for a character, avoid H3 Max in the panel, since it has no reference slots there.
Sume's docs note that clip limits are not uniform: wan-3.0 accepts 2 to 30 seconds and minimax-h3 accepts 5 to 15 seconds at native 480p or 768p, so the end frame also decides how long the clip can run.
- Need a landing frame and long runtime: Wan 3.0.
- Need a landing frame and stereo audio: MiniMax H3 or H3 Max.
- Need only a start frame and expressive motion: Grok Imagine.
- Not sure: Auto, which accepts an end frame and picks the serving model for you.
How should I prepare the two frames?
Keep both frames the same aspect ratio and the same subject scale, because the model has to invent the motion between them and a mismatch shows up as a morph or a cut. Use public HTTPS URLs for both images; the API takes URLs, not uploads.
Write the prompt about the movement, not the content of the frames. The frames already say what the subject looks like; the prompt should say how the camera or subject gets from one to the other, for example a slow push-in while the label turns toward the lens.
If the first result drifts, change the frames before changing the model. A second attempt with a closer end frame usually teaches you more than switching engines, and an Idempotency-Key per attempt keeps retries from creating duplicate jobs.
What does Sume not do here?
Sume does not give Kling 3.0 an end frame in the panel, and it does not offer Kling 4.0 keyframes. Kling says Kling 4.0 will officially launch in October; until Sume lists it in the catalog, treat it as not available here. For the Kling 4.0 comparison, see Kling 4.0 multiple keyframes versus Sume first and last frame.
Sources
Related posts
More in Media tools
- YouTube caption file for a Short: with timing or without timing?
YouTube's Upload file option asks for With timing or Without timing. Which to pick for a Short, and how Sume's transcript segments and burned-in captions fit.
- YouTube Shorts max 1080p: downscale a 4K vertical clip with Sume
YouTube's Shorts help says uploads have a maximum resolution of 1080p. Downscale a 2160x3840 clip to 1080x1920 with video-trim's exact-mode output conform.
- YouTube Shorts auto-captions not showing? Burn them in with Sume
YouTube says Shorts captions exist only in certain languages and on mobile devices. When that is not enough, Sume burns captions into the MP4 for $0.20.
- YouTube Shorts safe zone: place burned-in captions with anchor_ratio
YouTube's Shorts editor shows white lines and icons where overlays may be hidden. Move Sume's burned-in captions clear of them with placement.anchor_ratio.
Written by Sume