Product photo to video: first frame or reference for each SKU?

Use frame_images when the photo must open the clip, input_references when it should guide the look. Send both and Sume runs image-to-video.

5 min readSume
All posts

Send the product photo as a first frame when the clip must open on that exact image, and send it as a reference when the photo should guide the look of a new scene. On Sume's video API, frame_images starts image-to-video and input_references starts reference-to-video, and if you send both, frame_images wins.

That one choice decides whether your pack shot survives. A first frame keeps it. A reference borrows from it.

The two fields

The video generation docs describe both fields. frame_images entries carry a frame_type of first_frame or last_frame and start image-to-video. input_references are style or content references that the model uses as visual guidance, not as exact frames.

First frame versus reference on Sume's /v1/videos, read 2026-10-05
NeedFieldModeBehavior
Clip opens on the exact pack shotframe_images, first_frameimage-to-videoPhoto is the first frame
Clip starts and ends on stillsframe_images, first and lastimage-to-videoBoth frames are controlled where the model supports it
Product should appear in a new sceneinput_referencesreference-to-videoPhoto guides the look, not a frame
Both fields sentframe_images winsimage-to-videoinput_references are not used as the mode

Check the model first

Models differ in what they accept, so read supported_frame_images and supported_input_references for the model from GET /v1/videos/models before you send. The docs show that seedance-2 lists first_frame and last_frame, and that the Seedance 2.x models, Wan 3.0, MiniMax H3 and H3 Max accept audio and video references, while Gemini Omni Flash 1.1 accepts image and video references but no audio.

Reading the capability list first is cheaper than finding out from a failed job. A reference type that a model does not accept is a request you should not send, and on a sale night you want to know that before the batch starts.

Which SKU gets which mode

A pack shot on a white background is a first-frame case. You want the product to be recognizable in frame one, with motion after it: a slow orbit, a lid opening, confetti. A lifestyle scene, such as the same mug on a holiday table, is a reference case: the photo tells the model what the mug looks like, and the prompt describes the scene.

If you must keep the label legible, choose first frame, keep the motion small, and check the last second of the clip. Reference mode is more likely to restyle fine print.

A first-frame request

Here is a first-frame request for the Video Router with the 720p tier of Omni Flash 1.1, which is $0.125 per second at Sume's rate, so five seconds is $0.625.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sku-1042-open" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "image_url": "https://example.com/sku-1042.png",
    "prompt": "Slow orbit around the gift box, soft snowfall, no text",
    "resolution": "720p",
    "duration": 5,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

Using it in a batch

The Video Router lists three paths for the same model: image_to_video with image_url, reference_to_video with reference_image_urls of up to ten images, and text_to_video. In that route, you refer to references in the prompt as <IMAGE_REF_0> in list order. Native audio is always on for Omni Flash 1.1, so mute or replace it if the ad needs a different bed.

For a catalog of fifty SKUs, use first-frame mode for the pack shots, run one test clip per product category, and only then batch. A five-second clip costs $0.625 at 720p, so fifty clips are $31.25, and a bad prompt that you catch on the test saves you most of that.

Prepare the photos in three piles

Photo quality sets the ceiling for both modes. A first frame with a cluttered background gives you a clip that starts cluttered. A reference with a glare across the label gives the model a glare to reproduce. Before you queue a catalog, sort the photos into three piles: clean pack shots for first frame, lifestyle shots for reference, and photos that need a fix first.

The fix pile is where an image edit step pays off. A short edit pass that cleans a background or removes a reflection costs cents per image on Sume's image catalog, and it is much cheaper than discovering the problem in a rendered clip. Ideogram 4.5, for instance, is $0.075 for a medium edit. Then send the cleaned image as the first frame.

Keep your prompts short and concrete. Name one camera move, one lighting mood and one thing that must not change. 'Slow orbit, soft snowfall, label stays sharp' is a better prompt than a paragraph of adjectives, because you can check each clause against the result.

Finally, check the aspect ratio of the source photo against the ratio you ask for. A square pack shot in a 9:16 request will be extended or cropped by the model, and you may not like which. If the placement is vertical, shoot or edit the source to the vertical ratio first.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume