App screenshots to a 30-second product demo with Wan 3.0 on Sume
Turn six app screenshots into a 30 second demo on Sume: five 6 second first-and-last-frame Wan 3.0 clips cost $3.75 at 720p. Setup, limits, and what to check.

Six app screenshots become a 30 second demo on Sume with five Wan 3.0 clips of 6 seconds each, where every clip starts on one screenshot and ends on the next; at 720p the five clips cost $3.75 in total. Wan 3.0 supports an image-to-video start frame with an optional end frame, which is the field set this uses.
Why first and last frames
The idea is to let the model draw the movement between two real screens, while the first and last frames of each clip are your actual UI. That is more faithful than reference-to-video, where the Sume docs say the model uses images as visual guidance and not as exact frames.
What the request looks like
Sume's Video Router lists wan-3.0 as t2v, i2v with an end frame, and r2v, with audio, at 480p, 720p and 1080p and 2 to 30 seconds. A start and end image on one request is the i2v-plus-end mode. The code accepts image_url and end_image_url on the router, and frame_images with first_frame and last_frame on /v1/videos. The Sume schema also rejects mixing first or end frames with reference_*_urls, so choose one approach per request.
The cost
Six screens make five transitions, and five clips of 6 seconds make 30 seconds.
| Resolution | Per second | Five 6 s clips | Same plan, six 5 s clips |
|---|---|---|---|
| 480p | $0.0625 | $1.875 | $1.875 |
| 720p | $0.125 | $3.75 | $3.75 |
| 1080p | $0.25 | $7.50 | $7.50 |
Prepare the screens
Prepare the screens at the aspect ratio you will deliver. A product demo for desktop is 16:9; for a phone store listing, 9:16. Wan 3.0 takes adaptive, 16:9, 4:3, 1:1, 3:4 and 9:16, and a screenshot with a different shape will be fit to the request, so crop before you upload. Keep the text in your screens large and the number of elements low.
The honest limit
Models redraw pixels. Small interface text, numbers and logos can drift in the in-between frames, even when the first and last frame are exact. Treat the generated motion as a transition between faithful stills, keep each clip short, and overlay product copy in an editor if it must be exact. A screen recording is the right tool when the exact interface behavior is the point.
Submit the transitions
Send each pair as one job with its own idempotency key.
import asyncio, json, os, urllib.request
SCREENS = [f"https://example.com/screen{i}.png" for i in range(1, 7)]
def submit(i, a, b):
body = {"model": "wan-3.0", "resolution": "720p", "duration": 6,
"aspect_ratio": "16:9", "generate_audio": False,
"prompt": "Smooth screen-to-screen transition, clean UI, no new text",
"frame_images": [
{"type": "image_url", "image_url": {"url": a}, "frame_type": "first_frame"},
{"type": "image_url", "image_url": {"url": b}, "frame_type": "last_frame"}]}
req = urllib.request.Request("https://api.sume.com/v1/videos",
data=json.dumps(body).encode(), method="POST",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": f"demo-{i}-v1"})
return json.load(urllib.request.urlopen(req))["id"]
async def main():
pairs = list(zip(SCREENS, SCREENS[1:]))
print(await asyncio.gather(*(asyncio.to_thread(submit, i, a, b) for i, (a, b) in enumerate(pairs, 1))))
asyncio.run(main())Sound
Audio is switched off in the sample, since a demo normally carries your own music or voice-over. The code defaults generate_audio to true for Wan, so set it explicitly to match what you want, and check the field in the Video generation docs.
Pacing and retries
Order and pacing decide whether the demo reads. Pick the six screens that tell one story: sign-in, home, the key action, the result, a setting that matters and the confirmation. Six seconds between two screens is slow enough to follow a cursor or a swipe. If a transition feels long, make the same plan with six 5 second clips, which costs the same $3.75 at 720p.
Budget a retry per demo. If one transition in five distorts a button label, re-running that one clip costs 6 x $0.125 = $0.75 at 720p, not the whole $3.75. That is the main advantage of building the demo from separate clips.
Draft at 480p
Run the first pass at 480p. Five clips of 6 seconds are $1.875 at that tier, and you will see whether the model keeps your interface stable before you pay for 1080p, where the same plan is $7.50. If you plan to publish at 1080p, render only the clips you have already approved.
Versus one long prompt
Compared with one 30 second Wan clip from a prompt ($3.75 at 720p, the same price), the transition approach gives you control of every screen at the cost of five jobs and a join. Join the five MP4 files in your editor, add the voice-over, and review at full size. Vendor background: Alibaba lists image-to-video and reference-to-video among five pipelines in its Wan 3.0 repository (read 2026-10-05); the Sume limits are in the Video Router docs.
Sources
Related posts
More in Use cases
- Product demo video from a screen recording: $0.34 on Sume
Cut a 3-minute screen recording to a 30-second demo, add a voiceover and captions on Sume for about $0.34. Google Play autoplays only the first 30 seconds.
- Apple Podcasts 1.11: disclose AI voices in the audio and the metadata
Apple's podcast guidelines require prominent disclosure of synthetic voices and AI hosts, in the content and the metadata. What to write, and a record to keep.
- Approve the first frame before paying for motion: a $0.0094 draft gate
Review three $0.0094 draft stills, approve one, then pay $0.0835 for the final and $1.89 for the clip: $2.00 a SKU, against $5.75 for three blind clips.
- ArtStation 25 MB free vs 250 MB premium clips: bitrate budget
A 60-second ArtStation clip on a free account must average about 3.3 Mbps to fit in 25 MB. Premium's 250 MB allows about 33 Mbps. Table and a Python check.
Written by Sume