Turn a product changelog into a 30-second update video

Convert release notes into a short update video: pick three changes, write a 70-word script, voice it with TTS, add real screenshots and render with Timeline.

5 min readSume
All posts

To turn a changelog into a 30-second video, keep three changes, write about 70 words, generate the voiceover with Sume TTS, drop one real screenshot per change into a Timeline 1.0 render and let the voice length set the cut points. The machine cost is about $0.12: roughly $0.02 of TTS plus $0.10 for one rendered minute. You supply the screenshots.

Software teams ship notes faster than they ship video, so the useful habit is a template you re-run every release.

Pick three, not ten

Changelogs list everything. A 30-second video can carry three beats. As a worked example, take Raycast's v2.6 notes, published 30 September 2026: the host process uses about 50% less memory, extension processes use about 30% less memory at startup, and the Auto model has Lightweight, Balanced and Flagship preferences, with Balanced using about 70% fewer credits than Flagship. Those are three claims with numbers, which is what a short video can hold. Quote the vendor's wording exactly; do not round or embellish a number you read.

The pipeline

Steps and Sume list prices, from the Sume API pricing page and docs.sume.com, read 2026-10-01. Amounts are computed here.
StepEndpointCost
Voiceover, about 450 charactersPOST /v1/tts-router/generate with timestamps.words trueabout $0.02 at $0.0475 per 1,000 characters
Real screenshots inUpload via POST /v1/assets/upload-url, PUT, then POST /v1/assets/{id}/completeno generation charge listed
Assemble 30 secondsPOST /v1/timeline-1.0/render$0.10 (one rounded-up minute)
Total machine costabout $0.12

Cut on the words

Ask TTS for word timestamps and place each screenshot's start at the first word of its sentence. The completed TTS result carries monotonic words[] with start and end seconds, so you do not guess. Put the TTS file on audio.url with its measured duration, give each slot fit: "contain" or "blur" so a desktop screenshot is not cropped in a vertical frame, and add a 0.25-second fade between slots. If you want phone-sized text, burn a caption pass afterward at $0.20.

Use real screens, not generated ones

An image model can draw a plausible app window, but it will not match your product. Take screenshots from the product, or pull stills from a screen recording with POST /v1/video-frames, which is unbilled. Generated imagery belongs in the title card, not in the proof.

Limits

The script is yours. Sume does not verify that a changelog claim is true or still current, so link the vendor's notes in the post and re-read them on the day you publish. TTS lengths vary by voice, so measure the output, then trim the script rather than speeding the audio. A single TTS job takes up to 20,000 characters, far more than this needs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume