Ad from one product photo: a 4:5 still, a 9:16 clip and a 6 s cutdown
Three Sume calls turn one packshot into a 4:5 feed image, a 9:16 UGC-style clip and a short cutdown. Settings, order and what to check at each step.

With one product photo you can make a 4:5 feed image, a 9:16 UGC-style clip and a short cutdown in three calls: an image edit with openai/gpt-image-2.5 at aspect_ratio: "4:5", a Video Router clip that starts from that still, and a second, shorter render of the same prompt. Do them in that order, and look at each result before you pay for the next.
This is the shape of work a Meta or TikTok Shop seller does every week: one packshot, three placements. Sume ships the image and video calls; what it does not do is read the placement's own rules for you, so check the destination's current specs before you upload.
Call one: the 4:5 still
Send the packshot as input_references and describe the scene around the product. Keep quality at medium while you tune, then high. The Sume docs name 4:5 as Instagram portrait, 1080 by 1350. State that the product's label, colour and shape must not change, and name any text that must stay legible.
A request returns 200 with the image or 202 after the 30 second blocking budget; poll the job and store the Sume-hosted URL.
| Piece | Call | Settings |
|---|---|---|
| Feed still | POST /v1/images | gpt-image-2.5, 4:5, high |
| UGC-style clip | POST /v1/video-router/generate | seedance-2.5, 9:16, 720p, 4 to 30 s |
| Cutdown | Same call, shorter | Lower duration, same prompt |
| Safety | Idempotency-Key header | One per piece and version |
Calls two and three: the clip and the cutdown
seedance-2.5 accepts 4 to 30 seconds at 480p, 720p and 1080p, and Video Router bills the provider list price times 1.25 per output second. Write the prompt like a person filming on a phone, handheld, natural light, on a desk. Draft at 720p. Once the motion is right, keep the same prompt and cut duration for the shorter version, which is cheaper than re-thinking the idea.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: packshot-77-clip-12s-v1" \
-d '{
"model": "seedance-2.5",
"prompt": "A vertical UGC-style clip: hands pick up the bottle from a desk, turn it to show the label, set it down. Natural light, handheld.",
"resolution": "720p",
"duration": 12,
"aspect_ratio": "9:16",
"mode": "async"
}'Checks before upload
Watch the first two seconds with sound off, then read the label in a frame. If the label is wrong, generate the still again with the label named in the prompt rather than hoping the clip fixes it. Pull frames with POST /v1/video-frames if you want to compare the product against the packshot side by side.
A one-run alternative is a Format, which does the orchestration for you; the two-stage before-after post shows one. For the aspect behaviour on edits see auto versus omitted, and for doing this for many SKUs see the bulk post.
Mark generated creative as the destination requires; this post does not state those rules.
Keeping the three pieces consistent
The risk of making three assets separately is that the product drifts between them. Use the same packshot as the reference everywhere, repeat the same product description in each prompt, and do not rewrite the wording between calls except for the framing. If the clip's product looks different from the still's, send the approved still as the clip's first frame so the video starts from the product you already approved.
Tag each call with metadata where the endpoint allows it, and use an Idempotency-Key per piece, such as packshot-77-still-v1, packshot-77-clip-12s-v1 and packshot-77-cut-6s-v1. That naming gives you a clean log of what was made and lets a retry replay instead of double charging.
Finally, test the cutdown on its own. A 6 second version should not feel like the first 6 seconds of the long one; it needs its own hook in the first second. Rewrite the first line of the prompt for the short version and keep the rest.
A note on the cutdown
The cutdown is the piece most sellers skip, and it is often the one a feed rewards, since a short clip gets watched to the end more easily. Make it with the same model and a lower duration, and if the long clip has a payoff in the last seconds, say so in the short prompt so the motion starts toward it earlier. Review all three side by side before you schedule them.
Sources
Related posts
More in Use cases
- Medicare open enrollment video for insurance agents, with AI
Medicare open enrollment runs Oct 15 to Dec 7. Plan a short explainer series with an AI avatar, captions and one Timeline cut, with every claim checked by you.
- Meme and trend Shorts: when many creators upload the same thing
YouTube lists content uploaded many times by other creators as reused content. How to join a trend with something of your own, using trim and captions.
- Merchant Center video link dangerous products violation: fix it
Merchant Center disapproves a video_link as dangerous products: remove or replace the URL, wait 24-72 hours. The product stays live. Make a new clip with Sume.
- Meta Says 9:16 Video With Audio Cuts Cost Per Result 34.5%: Test It
Meta reports 34.5% lower cost per result for 9:16 video with audio and safe-zone messaging vs image ads. Build a fair test with Sume trim, captions and music.
Written by Sume