Five product photos to one 15-second ad: five clips and a Timeline

Animate five angles of one product with Wan 3.0 at $0.625 each, cut 3 s from each in a $0.10 Timeline render, add $0.20 captions: $3.43 for a 15 s vertical ad.

4 min readSume
All posts

To turn five photos of one product into a 15-second vertical ad, animate each photo as its own 5-second clip, then let a Timeline render keep 3 seconds of each and join them with short fades. On Sume at 720p with Wan 3.0 that costs 5 x $0.625 = $3.125 for the clips, $0.10 for the render and $0.20 for captions: $3.425 in total, read 2026-10-05 from the Sume pricing code and docs. Using Seedance 2 instead, the clips are 5 x $1.89 = $9.45 and the ad is $9.75.

Why five clips beat one long one

A single 15 s generation from one photo has to invent angles the photo does not show. Five photos give the model five real views, and the Timeline cut lets you keep only the best 3 seconds of each clip. Each clip uses the photo as first_frame on POST /v1/videos, as described in the videos docs. This estimate prices each Wan 3.0 clip at 5 s, so you pay for 5 s and use 3 s of it.

15 s ad from five photos, Sume prices read 2026-10-05
LineWan 3.0 720pSeedance 2 720p
Five 5 s clips$3.125$9.45
Timeline render, under 1 minute$0.10$0.10
Captions or price cue$0.20$0.20
Total$3.425$9.75

The Timeline document

Timeline 1.0 takes an audio spine plus ordered video[] slots, per the Timeline docs. With no voice-over, set audio.mode to "silence" and declare 15 seconds. Slots need increasing start values, a duration of at least 0.2 s, and the first slot cannot carry a transition. The helper below builds that document; call POST /v1/timeline-1.0/plan first, which is unbilled, to confirm segment count and cost.

import json


def timeline_body(clip_urls, seconds_each=3.0, fade=0.25):
    slots = []
    for i, url in enumerate(clip_urls):
        slot = {"source_url": url, "start": i * seconds_each,
                "duration": seconds_each, "fit": "cover"}
        if i:
            slot["transition"] = {"type": "fade", "duration": fade}
        slots.append(slot)
    return {"audio": {"mode": "silence",
                      "duration_seconds": seconds_each * len(clip_urls)},
            "video": slots}


if __name__ == "__main__":
    urls = ["https://media.sume.com/artifacts/demo/angle%d.mp4" % n for n in range(1, 6)]
    body = timeline_body(urls)
    print(body["audio"], len(body["video"]))
    print(json.dumps(body["video"][1]))

Checks before you pay for the render

Timeline imports only media.sume.com URLs, so each generated clip must already be a Sume artifact. The default output is 1080x1920, which suits a 9:16 placement. If a fade is longer than half the shorter neighbor, the plan rejects it, which is why the helper keeps fades at 0.25 s.

  • Order the clips from the hero angle to the detail shot, so the first cut shows the product whole.
  • Keep every photo on the same background so cuts read as one product, not five.
  • Render a plan call before a retake: only the clip that failed needs regenerating, not the whole set.

For a single-photo version of the same job, see one product photo to still, clip and cutdown; for the per-SKU multiplication at catalog scale, see the Seedance 2.5 per-SKU estimate.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume