Image-to-Video From a Product Still: End Frames and Grok Limits
Which video models turn one product photo into motion, which accept an end frame, and what Grok Imagine Video 1.5 requires on Sume.

A product photo is the safest first frame you have: the label, the color and the proportions are already right. What differs is what each model lets you do next. Some take only the first frame; some take a last frame so the clip can return to its start.
What each model takes
On Sume, kling-3, wan-3.0, minimax-h3, minimax-h3-max and gemini-omni-flash-1.1 accept a first frame plus an end frame. The Seedance family accepts first and last frames plus references. grok-imagine-video-1.5 is image-to-video only, requires image_url or first_frame_url, and has no end frame and no audio toggle; it runs 4 to 15 seconds at 480p or 720p for $0.0125 per second. xAI's own guide describes text, image and video-edit input, so Sume's row is narrower than the vendor's feature set.
| Model | First frame | End frame | Per second |
|---|---|---|---|
| grok-imagine-video-1.5 | Required | No | $0.0125 |
| wan-3.0 at 720p | Yes | Yes | $0.125 |
| minimax-h3-max at 768p | Yes | Yes | $0.10 |
| gemini-omni-flash-1.1 at 720p | Yes | Yes | $0.125 |
| kling-3, audio off | Yes | Yes | $0.14 |
A request that uses both frames
On POST /v1/videos, each entry in frame_images carries a frame_type of first_frame or last_frame. If you also send input_references, the frame images win and the request is treated as image-to-video. Google's Veo page says first and last frame are supported for Veo too, with the 8-second rule applying to 1080p and 4K.
Use the end frame for loops and for reveals where the final pose is the packshot. Skip it for open-ended motion, where forcing a destination often makes the middle of the clip stiff.
Budget approach
Draft on Grok at $0.0125 per second to check the framing and the motion direction: a 5-second test costs about 6 cents. Move to a model with an end frame only when the shot needs one. Remember that Sume exposes no seed on any video model, so a draft cannot be reproduced exactly at higher quality; the final is a new generation.
Sources
Related posts
More in Models
- Imagen 4 has no 4:5 ratio: get Instagram portrait from 3:4 on Sume
Imagen 4 Fast and Ultra on Sume list five ratios and none is 4:5. Generate 3:4 and trim 90 pixels, or pick a model that lists 4:5. A Pillow crop included.
- How many reference images does each Sume image model take?
Sume's image catalog lists 16 references for ChatGPT Image 2.5, 5 for Ideogram 4.5, 10 for most others and 0 for five text-only models. Full table.
- Irodori-TTS-v4-Large: Japanese cloning, 120 s reference, terms
Irodori-TTS-v4-Large is a 3.29B Japanese TTS model with emoji style control and Gemma terms. What the card says and how Sume's audio tools fit.
- Is FLUX.2 deprecated after FLUX 3? Keep production ids pinned on Sume
BFL's documentation says FLUX.2 remains fully supported for production image work. Pin the model id and read Sume's catalog before changing anything.
Written by Sume