Product turntable ad: which Sume video models take two frames
For a product turn, give a video model the front and back photo as first and last frames. Six Sume rows take both; Grok Imagine takes a first frame only.

For a product turntable ad, send a front photo as the first_frame and a back or side photo as the last_frame, and let the model fill the turn between them. On Sume's video catalog, gemini-omni-flash-1.1, wan-3.0, kling-3, minimax-h3, minimax-h3-max and the Seedance family accept an end frame; grok-imagine-video-1.5 does not, because it is an image-to-video row that takes a first frame only.
The catalog is the source here: Video Generation and the Video Router guide, both read on 2026-10-05. Pick by clip length, resolution and price, because every model that takes two frames does the same job in different limits.
Models that take a first and a last frame
The price column is the 720p or nearest native rate per second on Sume. Sume bills the provider list price times 1.25 and rounds the job up to the cent, so a one-second figure here is a rate, not a charge.
| Model id | Clip length | Resolutions | Sume rate per second |
|---|---|---|---|
| gemini-omni-flash-1.1 | 3 to 10 s | 360p, 720p, 1080p, 4K | $0.125 at 720p |
| wan-3.0 | 2 to 30 s | 480p, 720p, 1080p | $0.125 at 720p |
| kling-3 | 4 to 15 s | 720p, 1080p | $0.14 with audio off |
| minimax-h3 | 5 to 15 s | 480p, 768p | $0.075 at 768p |
| minimax-h3-max | 5 to 15 s | 480p, 768p, 1080p | $0.10 at 768p |
| seedance-2.5 | 4 to 30 s | 480p, 720p, 1080p | billed per 1,000 video tokens |
What a 5-second turn costs
A 5-second turn at 720p is $0.625 on gemini-omni-flash-1.1 or wan-3.0, $0.70 on kling-3 with audio off, and $0.375 on minimax-h3 at its native 768p. A turntable needs no sound, so generate_audio: false on Kling is the cheaper setting; Omni and H3 always make audio and you pay for it either way.
Draft first: Wan at 480p is $0.0625 a second, so a 5-second draft of the turn is about $0.32. Approve the angle pair, then pay for 1080p once.
Request
Two frame_images entries, one per frame_type. Both must be public HTTPS URLs.
import os, time, requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
payload = {
"model": "wan-3.0",
"prompt": "Slow 180 degree turntable of a ceramic mug on a white sweep, soft studio light",
"duration": 5,
"resolution": "480p",
"aspect_ratio": "1:1",
"frame_images": [
{"type": "image_url", "image_url": {"url": "https://example.com/mug-front.png"}, "frame_type": "first_frame"},
{"type": "image_url", "image_url": {"url": "https://example.com/mug-back.png"}, "frame_type": "last_frame"},
],
}
job = requests.post("https://api.sume.com/v1/videos", headers=H, json=payload).json()
while True:
time.sleep(30)
s = requests.get(job["polling_url"], headers=H).json()
if s["status"] in ("completed", "failed", "cancelled"):
break
print(s["status"], s.get("unsigned_urls"), s.get("usage"))Limits to plan around
Square output rules out Omni, which takes only 16:9 or 9:16. Kling takes 16:9, 9:16 and 1:1. No catalog model accepts seed, so a rerun gives a new take, not the same turn.
If the product has text on the label, check the last frame of the result first; the end frame is a target the model approximates, not an exact pixel match.
Sources
Related posts
More in Media tools
- Put a promo code on screen in an AI product video with caption cues
Burn a Black Friday code into a silent clip with $0.20 caption cues, key the retry by SKU and code, then pull stills with video-frames to check the text.
- Proof frame for a 1080x1920 export: video frames PNG at 1920
Pull a lossless 1080x1920 still from your vertical export with video frames (format png, max_edge 1920) to check captions and crop before you upload.
- Silent reference clip: reference_ingest's -60 LUFS gate and STT
reference_ingest marks a clip audio.silent at -60 LUFS or below and skips STT, so a silent reference costs no transcript. What to do next: plan new music.
- reference-ingest source_no_video_stream: audio-only file as input
source_no_video_stream means the reference file has no video track, such as an m4a. Send a clip with picture, or use audio detach to work with the sound.
Written by Sume