Same product in six Wan 3.0 shots: one reference set, one Timeline
Keep a product consistent across six Wan 3.0 clips on Sume: same reference images in every request, vary only the action line, then join in Timeline.

To keep one product looking the same across six Wan 3.0 shots, send the same set of reference_image_urls in every request, change only the action sentence of the prompt, and join the six clips with Timeline 1.0. Sume's Video generation docs say Wan 3.0 accepts image, video and audio references, and the catalog entry for wan-3.0 caps reference images at 10 per request. Alibaba's Wan 3.0 README (read 2026-10-05) lists up to 20 reference assets, so the model is built for multi-reference jobs; Sume's lower image cap is the one to plan around.
Build the reference set once
Pick three to five images: a straight-on hero, a three-quarter angle, a close-up of the label or detail, and optionally the product in use. Use clean backgrounds and the same lighting. Host them at stable https URLs and save the exact list. The list is your product bible: if you reorder it between clips, the outputs can differ.
- One hero image with the full product in frame.
- One angle that shows the back or side, so the model has seen the shape.
- One close-up of the part people recognize (cap, handle, label).
- Optional: one image of a hand holding it for scale.
- Do not mix product photos from different seasons or packaging versions.
- Name the files in order (01-hero, 02-side, 03-label) so the order you send is the order you saved.
Vary the action, hold the rest
Write one fixed sentence that describes the product and lighting, and paste it into every prompt. Then add a second sentence for the action of that shot. Six shots might be: slow turntable, hand lifts it, close-up of the label, pour or open, product on a shelf, final hero hold.
| Shot | Fixed line | Action line | Seconds |
|---|---|---|---|
| 1 | Same product sentence | Slow turntable on a plain surface | 4 |
| 2 | Same product sentence | A hand lifts it into frame | 4 |
| 3 | Same product sentence | Macro push-in on the label | 3 |
| 4 | Same product sentence | The lid opens, small steam or spray | 4 |
| 5 | Same product sentence | Set on a kitchen shelf, soft pan | 4 |
| 6 | Same product sentence | Final hero hold, locked camera | 3 |
Join them
After the six jobs finish, import the clips if they are not already media.sume.com artifacts, then call Timeline 1.0 with six ordered video[] slots, each with source_url, start and duration, and a fade or dissolve transition from the second slot on. The docs cap a transition at 1 second. POST /v1/timeline-1.0/plan runs the compile unbilled first, so you see the segment count and estimated cost before you render.
Mind the output size: the default render is 1080 by 1920, so a 16:9 clip is cropped with fit: cover unless you set output.width, output.height or fit yourself.
Set output.fps only if you know the rate you want. If you omit it, the docs say the render follows the rate of the sources, and a forced rate that differs from the clips repeats or drops frames and can cause judder on motion.
Where it still drifts
References steer the look; they do not lock it. Small print on a label can change between shots. If the label must be exact, take it out of the model's hands: render the clean product shots with Wan 3.0 and overlay a label still with Timeline compose where legibility matters.
Sources
Related posts
More in Use cases
- School bake sale poster to a 4-second clip: Seedance 2.0 Mini at 480p
The cheapest Seedance on Sume, seedance-2-mini, turns a bake-sale poster into a 4-second 480p clip for $0.35; Fast, 2.0 and 2.5 cost more.
- Search a podcast archive for a phrase and jump to the time: STT words
Transcribe episodes with Sume STT, keep words[] with start times in a small index, and find any phrase with a mm:ss link. Python, about 60 cents per hour.
- Swap the seasonal dish in a restaurant clip with Omni Flash edit
Keep one good restaurant clip and edit the dish for each holiday special with Gemini Omni Flash 1.1 video_to_video: $0.125 per output second at 720p.
- Seedance 2.5 takes 30 image references per pass: a shot-list ad
ByteDance lists 30 images, 10 clips and 10 audio files per Seedance 2.5 pass. How to turn a shot list into one seedance-2.5 request on Sume, with limits.
Written by Sume