Fashion lookbook video from flat-lay photos: 12 outfits, 9:16 cost

Animate 12 flat-lay outfit photos into 6-second 9:16 clips on Sume: $7.20 on H3 Max, $9.00 on Wan 3.0 or Omni, $41.64 on Seedance 2.5.

5 min readSume
All posts

To make a vertical lookbook from flat-lay photos, send each photo as the first frame of a 9:16 image-to-video request and describe a small camera move. On Sume a 6-second 9:16 clip at 720p costs $0.60 on MiniMax H3 Max at 768p and $0.75 on Wan 3.0 or Gemini Omni Flash 1.1. Twelve outfits are therefore $7.20 to $9.00, against $41.64 for Seedance 2.5.

The result is a short moving tile for each look, not a model wearing the clothes. If you want a person in the garment, that is a different job and needs a different input.

Twelve outfits, six seconds each, 9:16

All five models below accept 9:16. Sume bills the provider list price times 1.25 and rounds each clip up to the cent, so the twelve-clip total is twelve times the single price. Seedance 2.5 and the cheaper Seedance tiers are billed per video token, computed here at 720x1280.

6-second 9:16 image-to-video clips, Sume billable price (read 2026-10-07)
Model and resolutionOne clip12 outfits
MiniMax H3 Max, 768p$0.60$7.20
Wan 3.0, 720p$0.75$9.00
Gemini Omni Flash 1.1, 720p$0.75$9.00
Kling Video v3 Pro, sound off$0.84$10.08
Seedance 2 Mini, 720p$1.14$13.68
Seedance 2 Fast, 720p$1.82$21.84
Seedance 2.5, 720p$3.47$41.64

Shoot the flat-lay for the model, not for the shelf

  • Frame the photo at 9:16 before you upload; Omni Flash accepts only 16:9 and 9:16, and Kling only 16:9, 9:16 and 1:1.
  • Keep the garment inside the middle two thirds of the frame, because the move may crop the edges.
  • Use even light and a plain surface. Hard shadows get animated into odd shapes.
  • Photograph details separately, such as the buttons, the label and the fabric, instead of expecting a clip to invent them.

Prompt for a move, not a story

Fabric is easy for a video model to move and hard to keep identical. Ask for a gentle effect, such as a slow top-down push-in, a soft breeze lifting the sleeve, or a light sweep across the weave. Add what must remain fixed: "colours, pattern and stitching unchanged". Avoid asking for a person to appear, and avoid asking for the garment to be worn or folded during the clip. The model will invent one, and the garment will change to fit.

Six seconds is enough for a tile in a story. If you plan to put several outfits in one reel, render each tile separately and join them afterwards, so a failed render only costs you one tile.

import os
import requests

API = "https://api.sume.com/v1/videos"
HEADERS = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
PHOTOS = [
    "https://example.com/photo-1.jpg",
    "https://example.com/photo-2.jpg",
]

for url in PHOTOS:
    body = {
        "model": "wan-3.0",
        "prompt": "Slow top-down push-in, soft breeze lifts the sleeve, colours and stitching unchanged",
        "duration": 6,
        "resolution": "720p",
        "aspect_ratio": "9:16",
        "frame_images": [
            {
                "type": "image_url",
                "image_url": {"url": url},
                "frame_type": "first_frame",
            }
        ],
    }
    job = requests.post(API, headers=HEADERS, json=body, timeout=60)
    job.raise_for_status()
    print(url, job.json()["polling_url"])

Check the totals before a season launch

A 40-look collection on Wan 3.0 is $30.00 for one pass of 6-second clips. Budget two passes if you expect to redo one clip in four. The Sume job result returns usage.cost for every clip, so the first two or three renders tell you how far the real bill is from the plan.

Remember that Sume reserves the full amount from your balance when you submit, so queue a collection in batches that your balance can cover. A job that fails is refunded, and you can resubmit just that outfit.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume