Wan-Dancer 14B music-to-dance vs an audio reference on Wan 3.0

Wan-Dancer-14B (Apache 2.0, 2026-07-13) turns music into dance video locally. On Sume, Wan 3.0 accepts an audio reference on clips of 2-30 seconds.

4 min readSume
All posts

Wan-Dancer-14B is an Apache 2.0 music-to-dance model, announced July 13, 2026, whose paper title promises minute-scale coherent dance from music. On Sume the closest hosted route is wan-3.0, which the video docs list at 2-30 seconds and which honors audio references; it is a general video model, not a dance-specific one.

Wan-Dancer facts are from its model card, read 2026-10-01. Sume facts are from Video generation, read 2026-10-01.

What does Wan-Dancer release?

The card lists the license as apache-2.0, the pipeline as image-to-video, and says the model weights and inference code were released on July 13, 2026, with ComfyUI integration ticked. Its title is "A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation". The card excerpt gives install steps and downloads but no clip-length limit, so check the project page for the exact figure.

Wan-Dancer local model versus hosted Wan 3.0 on Sume, read 2026-10-01.
ItemWan-Dancer-14BSume `wan-3.0`
InputMusic, plus an image per the pipeline tagPrompt, plus references the catalog lists
AudioMusic is the driverAudio references are honored
LengthMinute-scale per the paper title2-30 seconds
Where it runsYour GPUHosted job

How does audio work on Sume?

Only models whose supported_input_references lists a type accept it. The docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. The catalog shows the advertised types as a list such as image_url, video_url and audio_url. An audio reference conditions the clip; it is not documented as beat-aligned choreography.

How do I cover a full song?

Most catalog models stop at 15 seconds, and wan-3.0 reaches 30. For a longer track, make clips and join them; the approach is in extend an AI video past 30 seconds. Expect visible seams between clips unless you chain frames.

When should I run Wan-Dancer locally instead?

When you need one continuous minute-scale dance driven by a full song and have the GPU for it. For short social clips with a reference track, a hosted wan-3.0 request avoids the setup. The dancer-avatar angle is covered in AI avatar hand gestures and dance moves.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume