Seedance prompt for sound: dialogue, ambience and music lines
Seedance generates voices, effects and music with the picture. How to write each as its own prompt line, what Sume's generate_audio flag does, and its limits.

Write Seedance's sound the way the model produces it: as separate layers. ByteDance describes Seedance 2.0 as generating background music, ambient sound effects and character voiceovers in parallel tracks, so a prompt with one line each for dialogue, ambience and music gives the model three clear jobs instead of one vague "with sound". On Sume the switch is generate_audio, and you get one mixed track back.
The line-by-line template here is a prompting habit, not an official ByteDance format. The vendor facts are only what the two pages below state; everything about wording is a test you should run yourself at 480p.
What do ByteDance and Dreamina say about Seedance audio?
The Seedance 2.0 launch post says the model uses dual-channel stereo and produces multi-track parallel output covering background music, ambient sound effects and character voiceovers. It also lists lip-synced multilingual dialogue as part of joint audio-video generation.
The Dreamina page for Seedance 2.5 adds two points: lip-sync for audio generations, and a claim that it removes random subtitles and unwanted background music, delivering clean output. That second point is why an explicit music line helps. If you do not ask for music, you should not assume you get it, and if you want none, say so.
| Fact | Source | What it means for the prompt |
|---|---|---|
| Dual-channel stereo audio | Seedance 2.0 launch post | Describe spatial sound only if you need it, such as a car passing left to right |
| Music, ambience and voiceover as parallel tracks | Seedance 2.0 launch post | Give each layer its own line |
| Lip-sync for audio generations | Dreamina Seedance 2.5 page | Quote spoken words and name the speaker |
| Removes random subtitles and unwanted music | Dreamina Seedance 2.5 page | State music or no music explicitly |
How do you write the three lines?
Start with the picture in one or two sentences, then add the layers. Dialogue goes in quotation marks and names who speaks. Ambience names the room or place and two or three sounds. Music names a genre and a role (underscore, beat, none), and says how loud it sits against the voice.
Keep the dialogue short. One sentence per speaker turn is easier to hold in sync, and a 4 to 8 second clip leaves little room for more. If a line matters legally or commercially, plan to caption it afterwards from your own text rather than trusting the generated speech word for word.
{
"model": "seedance-2.5",
"prompt": "A chef plates a dessert on a steel counter, locked medium shot. Dialogue: the chef says \"Last touch, the salt.\" Ambience: busy kitchen, a pan sizzling, distant plates. Music: soft acoustic underscore, quiet under the voice.",
"duration": 6,
"resolution": "480p",
"aspect_ratio": "16:9",
"generate_audio": true
}What does generate_audio do on Sume?
generate_audio is a boolean on the video generation request; omitted, it defaults to the model's audio capability. Sume returns a finished MP4 as a hosted artifact. The docs describe no separate stems or per-track volume, so voice, ambience and music arrive mixed. If you need to rebalance, detach the audio with the audio detach tool and work on it in a timeline, which is a separate step.
Seedance 2.x also accepts audio references: the docs say audio and video references are honored by the Seedance 2.x models, and the Video 1.0 page says reference audio needs at least one image or video reference alongside it. A reference sample is for steering, not a guarantee of an exact voice.
How do you iterate without overspending?
Change one layer at a time. Render at 480p, 4 to 6 seconds, keep the picture sentence identical, and alter only the ambience line, then only the music line. Listening to three short drafts costs less than one 1080p failure, and Sume holds the estimated cost at submit and refunds it if a job fails.
When the sound is right, rerun the winning prompt at the final resolution. If generated speech keeps missing, set generate_audio to false and add narration, or use a clean render and burn authored captions with the video captions API, which accepts your own cues and skips speech-to-text.
Sources
Related posts
More in Models
- Sora 2 snapshot ids removed Sept 24, 2026: the full OpenAI list
OpenAI's deprecations page lists five Sora ids removed from the API on Sept 24, 2026, with no replacement named. What to do with old ids and Sume's video list.
- Sume Auto video has no 1:1 or 21:9: which model to pin instead
Auto in Sume's Videos panel offers 720p or 1080p, 16:9 or 9:16 and 3 to 10 seconds. For square, ultrawide or longer clips, pin Wan 3.0, H3 or Kling 3.0.
- Reference limits per Sume video model: images, clips, audio
How many reference images, videos and audio clips each Sume video model takes on /v1/videos: Seedance 12 total, Wan 10/5/5, H3 9/3/3, Omni 10 and 3.
- Swap the product in a UGC clip with Gemini Omni Flash 1.1 video edit
Keep a UGC clip that works and change only the product: Sume's video router edit mode for gemini-omni-flash-1.1 takes a video_url and a one-line prompt.
Written by Sume