Turn a product changelog into a 30-second update video
Convert release notes into a short update video: pick three changes, write a 70-word script, voice it with TTS, add real screenshots and render with Timeline.

To turn a changelog into a 30-second video, keep three changes, write about 70 words, generate the voiceover with Sume TTS, drop one real screenshot per change into a Timeline 1.0 render and let the voice length set the cut points. The machine cost is about $0.12: roughly $0.02 of TTS plus $0.10 for one rendered minute. You supply the screenshots.
Software teams ship notes faster than they ship video, so the useful habit is a template you re-run every release.
Pick three, not ten
Changelogs list everything. A 30-second video can carry three beats. As a worked example, take Raycast's v2.6 notes, published 30 September 2026: the host process uses about 50% less memory, extension processes use about 30% less memory at startup, and the Auto model has Lightweight, Balanced and Flagship preferences, with Balanced using about 70% fewer credits than Flagship. Those are three claims with numbers, which is what a short video can hold. Quote the vendor's wording exactly; do not round or embellish a number you read.
The pipeline
| Step | Endpoint | Cost |
|---|---|---|
| Voiceover, about 450 characters | POST /v1/tts-router/generate with timestamps.words true | about $0.02 at $0.0475 per 1,000 characters |
| Real screenshots in | Upload via POST /v1/assets/upload-url, PUT, then POST /v1/assets/{id}/complete | no generation charge listed |
| Assemble 30 seconds | POST /v1/timeline-1.0/render | $0.10 (one rounded-up minute) |
| Total machine cost | about $0.12 |
Cut on the words
Ask TTS for word timestamps and place each screenshot's start at the first word of its sentence. The completed TTS result carries monotonic words[] with start and end seconds, so you do not guess. Put the TTS file on audio.url with its measured duration, give each slot fit: "contain" or "blur" so a desktop screenshot is not cropped in a vertical frame, and add a 0.25-second fade between slots. If you want phone-sized text, burn a caption pass afterward at $0.20.
Use real screens, not generated ones
An image model can draw a plausible app window, but it will not match your product. Take screenshots from the product, or pull stills from a screen recording with POST /v1/video-frames, which is unbilled. Generated imagery belongs in the title card, not in the proof.
Limits
The script is yours. Sume does not verify that a changelog claim is true or still current, so link the vendor's notes in the post and re-read them on the day you publish. TTS lengths vary by voice, so measure the output, then trim the script rather than speeding the audio. A single TTS job takes up to 20,000 characters, far more than this needs.
Sources
Related posts
More in Use cases
- Launch teaser from product shots: Gemini Omni Flash reference images
Feed up to 10 product references to gemini-omni-flash-1.1 on Sume, address them as IMAGE_REF_0 in the prompt, and get a 3-10 second teaser with native audio.
- Product photo on a white background: one image edit call on Sume
Turn a messy product photo into a clean white-background shot with one POST /v1/images edit. Which model, which fields, and what to check before you publish.
- Re-roll one shot in a stitched film without redoing the rest
Regenerate a single bad shot, swap its source_url in the Timeline 1.0 video array, and re-render the cut for $0.10 per output minute. The other shots stay.
- Listing photos to a narrated tour video with TTS and Timeline
Turn ten listing photos into a 45-second narrated tour: write the script, get word timestamps from TTS, then render stills on a Timeline spine for about $0.15.
Written by Sume