Release notes to video: three features, three shots, one Wan 3.0 clip

Turn a changelog into a 30 s Wan 3.0 video on Sume: one shot per feature, UI screenshots as references, exact labels burned as captions. Prompt and request.

5 min readSume
All posts

A release-notes video works best as one shot per feature. On Sume, send three screenshots of the new UI as reference_image_urls to wan-3.0, write a prompt with three numbered shots, ask for 24 to 30 seconds, and burn the feature names on as captions. Alibaba's Wan 3.0 README (read 2026-10-05) describes native 30 second generation and automatic scene splitting. Sume's Video Router docs show wan-3.0 accepting 2 to 30 seconds, so a three-feature update fits in one request.

Pick three, not ten

A changelog usually has more entries than a viewer can absorb. Choose the three that change what a customer does on Monday. Bug fixes and dependency bumps belong in the written notes. Each chosen item gets one sentence of the prompt and one reference image: the screen where the feature lives.

Crop the screenshots to the part of the screen that changed. A full-window capture with a toolbar, a sidebar and a feature panel gives the model many things to animate; a tight crop of the panel gives it one. Use the same aspect ratio for all three images so the cuts do not jump in framing, and avoid screenshots with personal data, since the model may reproduce faces or names it can read.

Prompt structure

Number the shots and give each one a single camera idea. Name the screen content in the prompt in plain words, but do not rely on the model to spell the feature title. The README claims text rendering in 12 languages; that is a reason to try it, not a reason to skip proofreading.

  • Shot 1 (0 to 9 s): slow push-in on the first screenshot, a cursor clicks the new button.
  • Shot 2 (9 to 18 s): hard cut to the second screenshot, a panel slides open.
  • Shot 3 (18 to 27 s): cut to the third screenshot, a badge appears, then hold.
  • Style line: flat product-demo look, no people, no logo animation.
  • Tail: ask for a 3 second hold on the last frame so you can overlay a call to action.

Request and what it costs to check

Use one request. The fields are the ones in the Video Router docs; reference_image_urls follows the same naming as the other reference fields on that page.

Then check the cuts. Sume's video inspect lets you ask for stills at named seconds with frames.at, up to 24 per call, so you can pull the frames at 8, 10, 17, 19 and 26 seconds and confirm the clip actually changed screens.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: relnotes-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "Three shots of a product update demo. Shot 1: push in on screen one, cursor clicks. Shot 2: cut to screen two, panel slides open. Shot 3: cut to screen three, badge appears, hold. Flat demo look, no people.",
    "reference_image_urls": ["https://example.com/feature1.png", "https://example.com/feature2.png", "https://example.com/feature3.png"],
    "resolution": "720p",
    "duration": 27,
    "aspect_ratio": "16:9",
    "mode": "async"
  }'

Add the exact labels

Burn the feature names with video captions cues, one per shot, timed to the cuts you found. If the model drew its own title text, cover it by making the cue large and placing it where the drawn text is not.

Then end with a still frame that carries the call to action. Sume's Timeline 1.0 can hold a still on a slot, so a branded end card can follow the generated clip without another model call.

Release-notes shots (read 2026-10-05)
ShotSecondsReferenceCaption cue
10 to 9Screen 1Feature name, 0.5 to 8.5
29 to 18Screen 2Feature name, 9.5 to 17.5
318 to 27Screen 3Feature name, 18.5 to 26.5

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume