Product photo to video: first frame or reference for each SKU?
Use frame_images when the photo must open the clip, input_references when it should guide the look. Send both and Sume runs image-to-video.

Send the product photo as a first frame when the clip must open on that exact image, and send it as a reference when the photo should guide the look of a new scene. On Sume's video API, frame_images starts image-to-video and input_references starts reference-to-video, and if you send both, frame_images wins.
That one choice decides whether your pack shot survives. A first frame keeps it. A reference borrows from it.
The two fields
The video generation docs describe both fields. frame_images entries carry a frame_type of first_frame or last_frame and start image-to-video. input_references are style or content references that the model uses as visual guidance, not as exact frames.
| Need | Field | Mode | Behavior |
|---|---|---|---|
| Clip opens on the exact pack shot | frame_images, first_frame | image-to-video | Photo is the first frame |
| Clip starts and ends on stills | frame_images, first and last | image-to-video | Both frames are controlled where the model supports it |
| Product should appear in a new scene | input_references | reference-to-video | Photo guides the look, not a frame |
| Both fields sent | frame_images wins | image-to-video | input_references are not used as the mode |
Check the model first
Models differ in what they accept, so read supported_frame_images and supported_input_references for the model from GET /v1/videos/models before you send. The docs show that seedance-2 lists first_frame and last_frame, and that the Seedance 2.x models, Wan 3.0, MiniMax H3 and H3 Max accept audio and video references, while Gemini Omni Flash 1.1 accepts image and video references but no audio.
Reading the capability list first is cheaper than finding out from a failed job. A reference type that a model does not accept is a request you should not send, and on a sale night you want to know that before the batch starts.
Which SKU gets which mode
A pack shot on a white background is a first-frame case. You want the product to be recognizable in frame one, with motion after it: a slow orbit, a lid opening, confetti. A lifestyle scene, such as the same mug on a holiday table, is a reference case: the photo tells the model what the mug looks like, and the prompt describes the scene.
If you must keep the label legible, choose first frame, keep the motion small, and check the last second of the clip. Reference mode is more likely to restyle fine print.
A first-frame request
Here is a first-frame request for the Video Router with the 720p tier of Omni Flash 1.1, which is $0.125 per second at Sume's rate, so five seconds is $0.625.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sku-1042-open" \
-d '{
"model": "gemini-omni-flash-1.1",
"image_url": "https://example.com/sku-1042.png",
"prompt": "Slow orbit around the gift box, soft snowfall, no text",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "9:16",
"mode": "async"
}'Using it in a batch
The Video Router lists three paths for the same model: image_to_video with image_url, reference_to_video with reference_image_urls of up to ten images, and text_to_video. In that route, you refer to references in the prompt as <IMAGE_REF_0> in list order. Native audio is always on for Omni Flash 1.1, so mute or replace it if the ad needs a different bed.
For a catalog of fifty SKUs, use first-frame mode for the pack shots, run one test clip per product category, and only then batch. A five-second clip costs $0.625 at 720p, so fifty clips are $31.25, and a bad prompt that you catch on the test saves you most of that.
Prepare the photos in three piles
Photo quality sets the ceiling for both modes. A first frame with a cluttered background gives you a clip that starts cluttered. A reference with a glare across the label gives the model a glare to reproduce. Before you queue a catalog, sort the photos into three piles: clean pack shots for first frame, lifestyle shots for reference, and photos that need a fix first.
The fix pile is where an image edit step pays off. A short edit pass that cleans a background or removes a reflection costs cents per image on Sume's image catalog, and it is much cheaper than discovering the problem in a rendered clip. Ideogram 4.5, for instance, is $0.075 for a medium edit. Then send the cleaned image as the first frame.
Keep your prompts short and concrete. Name one camera move, one lighting mood and one thing that must not change. 'Slow orbit, soft snowfall, label stays sharp' is a better prompt than a paragraph of adjectives, because you can check each clause against the result.
Finally, check the aspect ratio of the source photo against the ratio you ask for. A square pack shot in a 9:16 request will be extended or cropped by the model, and you may not like which. If the placement is vertical, shoot or edit the source to the vertical ratio first.
Sources
Related posts
More in Use cases
- Recall notice as a 9:16 and 16:9 avatar video: cost and captions
A recall notice needs to reach people on phones and on a site. Two 30-second Sume avatar jobs, one per aspect ratio, cost $14.70 at plus quality with captions.
- Proof watermark for client drafts of AI images, in Pillow
Stamp a tiled diagonal PROOF text over an image from Sume before sending drafts. Layer, rotate, alpha-composite, and why it is a courtesy and not security.
- Proposal walkthrough video with an AI avatar: 40 seconds, by tier
Turn a sales quote into a 40-second avatar walkthrough: scope, timeline and price in three scenes, the request body, and the per-quote cost on each Sume tier.
- Quarterly business review recap as an avatar video, three scenes
A QBR recap fits one 45-second Sume avatar video: result, one risk, next steps. A Plus render is $11.03; keep exact figures in the email, not the clip.
Written by Sume