Fashion lookbook video from flat-lay photos: 12 outfits, 9:16 cost
Animate 12 flat-lay outfit photos into 6-second 9:16 clips on Sume: $7.20 on H3 Max, $9.00 on Wan 3.0 or Omni, $41.64 on Seedance 2.5.

To make a vertical lookbook from flat-lay photos, send each photo as the first frame of a 9:16 image-to-video request and describe a small camera move. On Sume a 6-second 9:16 clip at 720p costs $0.60 on MiniMax H3 Max at 768p and $0.75 on Wan 3.0 or Gemini Omni Flash 1.1. Twelve outfits are therefore $7.20 to $9.00, against $41.64 for Seedance 2.5.
The result is a short moving tile for each look, not a model wearing the clothes. If you want a person in the garment, that is a different job and needs a different input.
Twelve outfits, six seconds each, 9:16
All five models below accept 9:16. Sume bills the provider list price times 1.25 and rounds each clip up to the cent, so the twelve-clip total is twelve times the single price. Seedance 2.5 and the cheaper Seedance tiers are billed per video token, computed here at 720x1280.
| Model and resolution | One clip | 12 outfits |
|---|---|---|
| MiniMax H3 Max, 768p | $0.60 | $7.20 |
| Wan 3.0, 720p | $0.75 | $9.00 |
| Gemini Omni Flash 1.1, 720p | $0.75 | $9.00 |
| Kling Video v3 Pro, sound off | $0.84 | $10.08 |
| Seedance 2 Mini, 720p | $1.14 | $13.68 |
| Seedance 2 Fast, 720p | $1.82 | $21.84 |
| Seedance 2.5, 720p | $3.47 | $41.64 |
Shoot the flat-lay for the model, not for the shelf
- Frame the photo at 9:16 before you upload; Omni Flash accepts only 16:9 and 9:16, and Kling only 16:9, 9:16 and 1:1.
- Keep the garment inside the middle two thirds of the frame, because the move may crop the edges.
- Use even light and a plain surface. Hard shadows get animated into odd shapes.
- Photograph details separately, such as the buttons, the label and the fabric, instead of expecting a clip to invent them.
Prompt for a move, not a story
Fabric is easy for a video model to move and hard to keep identical. Ask for a gentle effect, such as a slow top-down push-in, a soft breeze lifting the sleeve, or a light sweep across the weave. Add what must remain fixed: "colours, pattern and stitching unchanged". Avoid asking for a person to appear, and avoid asking for the garment to be worn or folded during the clip. The model will invent one, and the garment will change to fit.
Six seconds is enough for a tile in a story. If you plan to put several outfits in one reel, render each tile separately and join them afterwards, so a failed render only costs you one tile.
import os
import requests
API = "https://api.sume.com/v1/videos"
HEADERS = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
PHOTOS = [
"https://example.com/photo-1.jpg",
"https://example.com/photo-2.jpg",
]
for url in PHOTOS:
body = {
"model": "wan-3.0",
"prompt": "Slow top-down push-in, soft breeze lifts the sleeve, colours and stitching unchanged",
"duration": 6,
"resolution": "720p",
"aspect_ratio": "9:16",
"frame_images": [
{
"type": "image_url",
"image_url": {"url": url},
"frame_type": "first_frame",
}
],
}
job = requests.post(API, headers=HEADERS, json=body, timeout=60)
job.raise_for_status()
print(url, job.json()["polling_url"])
Check the totals before a season launch
A 40-look collection on Wan 3.0 is $30.00 for one pass of 6-second clips. Budget two passes if you expect to redo one clip in four. The Sume job result returns usage.cost for every clip, so the first two or three renders tell you how far the real bill is from the plan.
Remember that Sume reserves the full amount from your balance when you submit, so queue a collection in batches that your balance can cover. A job that fails is refunded, and you can resubmit just that outfit.
Sources
- fal: Wan 3.0 image-to-video (read 2026-10-07)
- fal: Gemini Omni Flash 1.1 image-to-video (read 2026-10-07)
- fal: MiniMax H3 Max text-to-video (read 2026-10-07)
- fal: Seedance 2.5 image-to-video (read 2026-10-07)
- fal: Kling Video v3 Pro image-to-video (read 2026-10-07)
- Sume docs: Video generation
- Sume docs: Video Router
Related posts
More in Use cases
- Film noir black-and-white AI video prompt for Gemini Omni
A black-and-white noir prompt for Gemini Omni Flash: lighting words, a no-dialogue sound line and a 10-second Sume request with prices for 720p and 1080p.
- What does a first-time home buyer video series cost with an avatar?
Six 30-second avatar tips cost $33.12 at standard, $44.10 at plus or $99.00 at max on Sume, plus $0.95 once for the avatar. Rates and script limits.
- Foggy forest morning clip with ambient sound: Gemini Omni, 10s
Prompt a 10-second foggy forest mood clip in Gemini Omni with ambient birds and drips. Sume request, why people-free scenes are easier, and cost per resolution.
- Food truck weekly special clip: Gemini Omni Flash 1.1, vertical, sound
Make a 3 to 10 second vertical clip of this week's special with Gemini Omni Flash 1.1 on Sume: native synced audio, 9:16, from a dish photo as a reference.
Written by Sume