Audio-reactive music visualizer: does Sume render waveforms?
No waveform or spectrum bars in Sume's docs. What ships: a still held over your audio on a timeline, or a video model that takes audio as a reference.

Sume does not document an audio-reactive visualizer, meaning bars, waveforms or a spectrum that move with the sound. The docs I read describe three things instead: Timeline 1.0 holds a still for as long as your audio spine runs ($0.10 per output minute), video generation models that accept audio as a reference input, and a video filter allowlist made of pixel filters. If you need real waveform bars, draw them elsewhere and import the finished video.
What does the timeline do with a cover and a song?
Timeline 1.0 takes an audio spine of up to 1800 seconds and ordered slots, and a still is accepted as a static hold. A track with cover art becomes one MP4 with the art on screen for the full song, at the 1080 by 1920 default or any even width and height from 256 to 2160. It does not animate the art with the sound. Add a cross-fade between two or three covers if you want some change, using fade, dissolve or the wipe types.
Which generation models take audio as a reference?
The Video generation page states that the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references, while Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio. An audio reference is an input the model can use. The page does not promise that motion lands on your beat, so judge each result by eye and listen for sync before you publish.
| Option | What it does | What it does not do |
|---|---|---|
| Timeline still plus spine | Cover art over the full track, $0.10 per minute | Animate with the sound |
| Video model with audio reference | Generates motion with the audio as an input | Documented beat-sync |
| Video filter | Pixel filters on one clip (tone, blur, geometry, fade) | Draw bars from audio |
| Music Router | Makes the track, $0.125 fixed | Make the visuals |
What about the filter graph?
Video filter accepts a filters-only ffmpeg graph from an allowlist the docs describe as tone, blur, geometry, fade and internal compositing. It works on one clip and returns a new MP4 at $0.02. The docs list trim, setpts, drawtext, subtitles, movie and lut3d as not allowed, and the graph cannot read files. A graph that draws from audio is not part of that surface.
What would a practical workflow be?
Make the track in Music Router or bring your own, generate or import the cover, and render a still-over-audio timeline for the release. For a livelier teaser, generate a short clip with a model that accepts audio references, trim the best seconds with video trim ($0.02), and put it in the timeline slot ahead of the cover. Check the catalog for the model's current limits first.
Sources
Related posts
More in Use cases
- Christmas carol sing-along video: lyric lines as timed cues
Trim your choir's recording, then burn each lyric line as a timed caption cue so a congregation can sing along. Two jobs, $0.22 at listed rates.
- Course Black Friday offer clip: lesson excerpt under a price card
Trim 20 seconds of a real lesson, stack a price card above it with timeline compose, and let captions transcribe the lesson. Three jobs, $0.24 listed.
- DJ set teaser: flyer above the booth clip, lineup as captions
Stack a night's flyer over 15 seconds of booth footage with timeline compose ($0.02), then burn set times as cues ($0.20). The clip's own audio stays.
- Double-exposure portrait from two photos with Nano Banana 2.1
Blend a portrait and a landscape into one double exposure: two input_references, a pinned 4:5 ratio, Nano Banana 2.1 at 1K, and the 10-reference cap on Sume.
Written by Sume