Wan-Dancer 14B music-to-dance vs an audio reference on Wan 3.0
Wan-Dancer-14B (Apache 2.0, 2026-07-13) turns music into dance video locally. On Sume, Wan 3.0 accepts an audio reference on clips of 2-30 seconds.

Wan-Dancer-14B is an Apache 2.0 music-to-dance model, announced July 13, 2026, whose paper title promises minute-scale coherent dance from music. On Sume the closest hosted route is wan-3.0, which the video docs list at 2-30 seconds and which honors audio references; it is a general video model, not a dance-specific one.
Wan-Dancer facts are from its model card, read 2026-10-01. Sume facts are from Video generation, read 2026-10-01.
What does Wan-Dancer release?
The card lists the license as apache-2.0, the pipeline as image-to-video, and says the model weights and inference code were released on July 13, 2026, with ComfyUI integration ticked. Its title is "A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation". The card excerpt gives install steps and downloads but no clip-length limit, so check the project page for the exact figure.
| Item | Wan-Dancer-14B | Sume `wan-3.0` |
|---|---|---|
| Input | Music, plus an image per the pipeline tag | Prompt, plus references the catalog lists |
| Audio | Music is the driver | Audio references are honored |
| Length | Minute-scale per the paper title | 2-30 seconds |
| Where it runs | Your GPU | Hosted job |
How does audio work on Sume?
Only models whose supported_input_references lists a type accept it. The docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. The catalog shows the advertised types as a list such as image_url, video_url and audio_url. An audio reference conditions the clip; it is not documented as beat-aligned choreography.
How do I cover a full song?
Most catalog models stop at 15 seconds, and wan-3.0 reaches 30. For a longer track, make clips and join them; the approach is in extend an AI video past 30 seconds. Expect visible seams between clips unless you chain frames.
When should I run Wan-Dancer locally instead?
When you need one continuous minute-scale dance driven by a full song and have the GPU for it. For short social clips with a reference track, a hosted wan-3.0 request avoids the setup. The dancer-avatar angle is covered in AI avatar hand gestures and dance moves.
Sources
Related posts
More in Use cases
- Apple Watch Ultra 4 screenshot size: 422x514 from Sume
Apple Watch Ultra 4 screenshots are 422 x 514 pixels. Sume's image API cannot output that size, but a Timeline still plus a PNG frame can.
- Wix Video minimum size 480x470: crop and scale with video filter
Wix Video accepts 480x470 to 1920x1080. Sume video-filter crops by fractions and runs scale or pad in a filtergraph, with a free check before the encode.
- Wix video poster 1944x2880: stills from a clip with video frames
Wix lists a video poster at 1944 x 2880 px. Sume video-frames makes JPG or PNG stills from a clip, but max_edge stops at 2160, so it cannot upscale to 2880.
- Wix Video trailer size: cut a short trailer with video trim
Wix lists trailers at up to 100 MB. Sume video trim cuts an exact-length clip with duration and output width, height and fps. It has no file-size setting.
Written by Sume