Music clip with Seedance 2.5: audio reference plus performer image
A Seedance 2.5 request on Sume takes reference_audio_urls only with a reference image or video. Build a 15 to 30 s music-driven clip and see the cost.

To build a music-driven clip with Seedance 2.5 on Sume, send your track in reference_audio_urls together with at least one reference_image_urls entry for the performer or scene. Audio alone is rejected with a 400. Set duration to the length of the section you want, from 4 to 30 seconds, and keep the audio file itself short and clean.
What the request needs
The Sume schema read on 2026-10-04 says reference_audio_urls requires at least one reference image or video. It also says first and last frame fields cannot be combined with reference_*_urls, so a music clip built from references starts from the images rather than from a fixed opening frame. On Seedance models Sume allows up to 3 audio references and 12 references in total.
BytePlus says Seedance 2.5 accepts audio files as references and lists up to 10 on its own platform (BytePlus: What is Seedance 2.5). Sume's lower cap applies to requests through Sume.
| Field | Role | Rule on Sume |
|---|---|---|
| reference_audio_urls | the track | 1 to 3 files; needs an image or video reference |
| reference_image_urls | performer, set, outfit | up to 9 images |
| duration | section length | whole seconds, 4 to 30 |
| image_url / end_image_url | opening or closing frame | cannot be mixed with reference fields |
Write the prompt around the music
Say what the performer does and how the camera moves, then describe the tempo in plain words: slow sway, steady walk, fast cuts. Do not paste lyrics into the prompt expecting exact lip-sync; Sume's docs do not promise it, and you should check the result yourself.
If you need guaranteed words on screen, add them afterwards with authored caption cues rather than relying on the generated frames.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: music-clip-001" \
-d '{
"model": "seedance-2.5",
"prompt": "The singer from the image walks along a rooftop at dusk, slow camera orbit, moving with the rhythm of the track.",
"reference_image_urls": ["https://example.com/singer.png"],
"reference_audio_urls": ["https://example.com/chorus.mp3"],
"resolution": "720p",
"duration": 15,
"aspect_ratio": "9:16",
"mode": "async"
}'Cost of a music-driven clip
Sume prices Seedance by output tokens, so the length and resolution set the price. The Video Router docs state a 1.25 margin over list. A full verse at 720p is the practical sweet spot between a throwaway draft and a final.
| Length | Price |
|---|---|
| 4 s | $2.31 |
| 10 s | $5.78 |
| 15 s | $8.67 |
| 30 s | $17.33 |
Where this fits
A 15 to 30 second clip suits a teaser or a looping social cut, not a full song. For longer pieces, generate several clips and join them on a timeline with the track as the audio spine; the Timeline 1.0 docs describe one audio spine plus ordered video slots.
Sources
Related posts
More in Use cases
- Muted product loop in four languages: burn captions from cues
A silent clip fails speech-to-captions as caption_no_speech. Pass cues with text, start and end to burn four translations at $0.20 a job.
- Naver Clip video: a 9:16 Sume clip with Korean captions
Making a vertical 9:16 clip for Naver Clip? Generate it on Sume, then burn Korean captions with the korean-ad style and language ko.
- Near-duplicate AI images: perceptual-hash dedupe before human review
Four-up image batches often contain near twins. Use a 64-bit difference hash in Pillow to drop duplicates before a person reviews them. Python for Sume results.
- Neighborhood holiday lights clip for agents from three street photos
A real estate agent's neighborhood lights tour from three of your own street photos: three 4-second Wan 3.0 shots at 480p and a join, about $0.85 on Sume.
Written by Sume