30-second lyric or music clip with an audio reference: $8.07 vs $1.88
Both seedance-2.5 and wan-3.0 take an audio reference on Sume. A 30-second clip is $8.07 at 480p on Seedance 2.5 and $1.88 at 480p on Wan 3.0.

Yes: seedance-2.5 and wan-3.0 both accept an audio reference on Sume, and both reach 30 seconds in one job. At 480p a 30-second clip is $8.07 on Seedance 2.5 and $1.88 on Wan 3.0. That is the whole decision for most music clips: Wan is the cheap way to try a track, and Seedance is the one to move to when a take needs more control.
Sume's video docs list Seedance 2.x, Wan 3.0, MiniMax H3 and MiniMax H3 Max as the models that accept audio and video references. Gemini Omni Flash 1.1 accepts video references but not audio (video generation docs).
How to pass the track
Send the audio as an input_references entry of type audio_url with a public HTTPS URL. Add image references for the look, and put the sections of the song in the prompt as beats.
{
"model": "wan-3.0",
"prompt": "0-10s: slow dolly through a neon alley. 10-20s: close on a singer's hands on a synth. 20-30s: wide shot, crowd lights rise.",
"duration": 30,
"resolution": "480p",
"input_references": [
{"type": "audio_url", "audio_url": {"url": "https://example.com/track.mp3"}},
{"type": "image_url", "image_url": {"url": "https://example.com/look.png"}}
]
}Price for the full 30 seconds
Wan 3.0 lists at $0.05, $0.10 and $0.20 per second; Sume bills list times 1.25 and rounds up to the cent. Seedance 2.5 is priced per video token and Sume bills that times 1.25 as well (video router docs).
| Model | 480p | 720p | 1080p |
|---|---|---|---|
| wan-3.0 | $1.88 | $3.75 | $7.50 |
| seedance-2.5 | $8.07 | $17.34 | $42.65 |
What the reference does and does not do
A reference is guidance. Do not expect frame-accurate lip-sync to a sung vocal from either id. If you need a performer's mouth to match a track, look at Sume's lip-sync and avatar options instead of a general video model.
Rights matter here: only send audio you own or have licensed, and keep the license note next to the output.
A cheap way to find the cut
Run the 30-second Wan 3.0 draft at 480p for $1.88. If the pacing works, you have the beat map. Then re-run on Seedance 2.5 only if you need its reference handling, or on Wan at 720p for $3.75. Do not start at Seedance 1080p: one miss is $42.65.
Practical notes
For longer songs, split the track into 30-second sections and generate each as its own job with the matching audio excerpt. That keeps each job at the 30-second ceiling and lets you redo a single section. A three-minute song is six jobs, about $11.28 on Wan 3.0 at 480p.
Finally, join the sections on a timeline and mix the full original track under them. The audio reference steers the generation; it is not a substitute for the master.
Sources
Related posts
More in Comparisons
- MAI-Transcribe-2-Streaming partials vs Sume STT terminal webhook
MAI-Transcribe-2-Streaming sends intermediate and final results as audio streams in. Sume STT sends one signed terminal callback. How to design for each.
- MAI custom voice SSML (ttsembedding) vs Sume voice ids: what is gated
Microsoft's MAI-Voice-2.1 custom voice sits behind Limited Access Review and uses a speakerProfileId in SSML. Sume picks voices by id, with no clone widget.
- Edit models Morphic names versus what the Sume catalog lists
Morphic runs image edits on Seedream 5.0 Pro, Nano Banana 2 and GPT Image 2.5 until Ideogram 4.5 arrives. Which of these Sume lists, and where to check.
- AI music cost per track: Sume Music vs Lyria 3.5 and ACE Step
Sume Music is a flat $0.125 per generation ($12.50 per 100). Google lists Lyria 3.5 at $0.08 per song ($8.00); fal lists ACE Step at $0.0002 per second.
Written by Sume