Product video for ecommerce: a white-background photo to 9:16
Turn a white-background product photo into a 5-second vertical scene with POST /v1/videos, seedance-2 and an image_url input reference.

To make an ecommerce product video from a plain white-background photo, call POST /v1/videos with model: "seedance-2", the photo in input_references, aspect_ratio: "9:16" and a short duration. Seedance 2.0 lists 480p, 720p and 1080p, durations from 4 to 15 seconds, and image, video and audio references in Sume's catalog. A reference conditions the whole clip on your item, which is what you want when the product has to look like the one you ship.
Source: Video generation: the /v1/videos API. Catalog values can change, so read GET /v1/videos/models for the current ones before you hard-code anything.
Reference or first frame?
The API infers the mode from which field you send. input_references makes it reference-to-video: the image conditions the full clip and the model draws the scene around it. frame_images with frame_type: "first_frame" makes it image-to-video: the clip opens on exactly that frame. If you send both, frame_images wins. For a white-background packshot, references are usually the better fit, because a first frame of a floating product on white is a poor opening for a feed video.
The request
Describe the scene and the motion, and keep the product description short; the image carries the product. A durable result needs a poll: the submit returns 202 with an id, then GET /v1/videos/{id} until completed, then GET /v1/videos/{id}/content. Add callback_url if you would rather be called. Do not send seed or size: the API rejects both in v1.
import json, os, time, urllib.request
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json"}
def call(method, url, body=None):
data = json.dumps(body).encode() if body else None
req = urllib.request.Request(url, data=data, method=method, headers=H)
with urllib.request.urlopen(req) as r:
return json.load(r)
job = call("POST", "https://api.sume.com/v1/videos", {
"model": "seedance-2",
"prompt": "The bottle on a sunlit bathroom shelf, slow orbit, soft steam, "
"shallow depth of field",
"input_references": [
{"type": "image_url",
"image_url": {"url": "https://media.sume.com/img/demo/bottle-white.png"}}
],
"aspect_ratio": "9:16",
"resolution": "720p",
"duration": 5,
})
while job["status"] not in ("completed", "failed", "cancelled"):
time.sleep(5)
job = call("GET", "https://api.sume.com/v1/videos/" + job["id"])
print(job["status"])
Checking that it is still your product
Pull stills from the result with POST /v1/video-frames and compare them to the photo. That endpoint takes a video_url on media.sume.com and either at[] times or an fps, always answers 202, and is billed by its own Modal compute rather than as a generation, so import or pass a hosted copy of the clip first. Compare colour, label text and proportions: colour, label text and proportions. Generated motion tends to drift in the details that identify a product, especially text on a label. If the label drifts, shorten the clip, simplify the motion in the prompt, or keep the label area out of the camera's path. The stored post on products looking the wrong size in hand covers scale.
| Field | Value | Note |
|---|---|---|
| model | seedance-2 | Read the catalog for current values |
| duration | 4 to 15 seconds | Whole seconds |
| resolution | 480p, 720p, 1080p | Pick 720p to test |
| aspect_ratio | 9:16 among others | 21:9 to 9:16 listed |
| input_references | image_url, video_url, audio_url | Image used here |
Status values and cost
The submit answers 202 with a pending status. The job then moves through in_progress to completed, failed or cancelled, so the poll loop above stops on all three terminal values; a loop that only checks for completed and failed would spin forever on a cancelled job. The finished job lists unsigned_urls that point at /v1/videos/{id}/content, and the same job can be read at GET /v1/jobs/{id}/status and /result in the normal Sume envelope.
The completed job also carries a usage.cost figure, so you can log what each SKU cost as you go. Sume bills the provider list rate times 1.25, which is why a cheap resolution like 720p is the right choice for the first test of a new product, and 1080p belongs to the final pass (Video generation: the /v1/videos API).
A last practical point: the API rejects seed and size in v1 with 400 unsupported_parameter, so you cannot pin a seed to get an identical second take. If a clip is close but not right, change the prompt or the reference and generate again, and keep the ids of the takes you rejected so the notes on why are attached to something.
When to scale this up
Test one SKU first, look at the result, then queue the rest with one stable key per SKU. A catalogue run is a loop over products with the same prompt template, not a different prompt for each. Keep the clean white-background photo as the master so a re-run starts from the same input.
Sources
Related posts
More in Use cases
- Real-estate agent, 25 listings a month: staged photos and voiceover
25 listings with six AI-edited photos, a voiceover and a music bed cost about $14.60 a month on Sume. Unit prices and the sum.
- Reddit 15-second ad: join hook, demo and offer, then plan first
Build a 15-second ad from three clips on Timeline 1.0 and check its length and cost with the unbilled /plan call before you pay for a render.
- Reddit 15-second engaged views: does your voiceover CTA land?
Reddit's Engaged Video Views beta bills videos over 15 s at 15 s. Use TTS word timestamps to check that your call to action is spoken before second 15.
- Reddit 15-second video ad: burn captions so it reads on mute
For a 15-second Reddit video ad, burn captions into the cut with POST /v1/video-captions: one $0.20 job, style slam for English, no new edit.
Written by Sume