Seedance prompt for sound: dialogue, ambience and music lines

Seedance generates voices, effects and music with the picture. How to write each as its own prompt line, what Sume's generate_audio flag does, and its limits.

5 min readSume
All posts

Write Seedance's sound the way the model produces it: as separate layers. ByteDance describes Seedance 2.0 as generating background music, ambient sound effects and character voiceovers in parallel tracks, so a prompt with one line each for dialogue, ambience and music gives the model three clear jobs instead of one vague "with sound". On Sume the switch is generate_audio, and you get one mixed track back.

The line-by-line template here is a prompting habit, not an official ByteDance format. The vendor facts are only what the two pages below state; everything about wording is a test you should run yourself at 480p.

What do ByteDance and Dreamina say about Seedance audio?

The Seedance 2.0 launch post says the model uses dual-channel stereo and produces multi-track parallel output covering background music, ambient sound effects and character voiceovers. It also lists lip-synced multilingual dialogue as part of joint audio-video generation.

The Dreamina page for Seedance 2.5 adds two points: lip-sync for audio generations, and a claim that it removes random subtitles and unwanted background music, delivering clean output. That second point is why an explicit music line helps. If you do not ask for music, you should not assume you get it, and if you want none, say so.

Audio facts from vendor pages (read 2026-10-02)
FactSourceWhat it means for the prompt
Dual-channel stereo audioSeedance 2.0 launch postDescribe spatial sound only if you need it, such as a car passing left to right
Music, ambience and voiceover as parallel tracksSeedance 2.0 launch postGive each layer its own line
Lip-sync for audio generationsDreamina Seedance 2.5 pageQuote spoken words and name the speaker
Removes random subtitles and unwanted musicDreamina Seedance 2.5 pageState music or no music explicitly

How do you write the three lines?

Start with the picture in one or two sentences, then add the layers. Dialogue goes in quotation marks and names who speaks. Ambience names the room or place and two or three sounds. Music names a genre and a role (underscore, beat, none), and says how loud it sits against the voice.

Keep the dialogue short. One sentence per speaker turn is easier to hold in sync, and a 4 to 8 second clip leaves little room for more. If a line matters legally or commercially, plan to caption it afterwards from your own text rather than trusting the generated speech word for word.

{
  "model": "seedance-2.5",
  "prompt": "A chef plates a dessert on a steel counter, locked medium shot. Dialogue: the chef says \"Last touch, the salt.\" Ambience: busy kitchen, a pan sizzling, distant plates. Music: soft acoustic underscore, quiet under the voice.",
  "duration": 6,
  "resolution": "480p",
  "aspect_ratio": "16:9",
  "generate_audio": true
}

What does generate_audio do on Sume?

generate_audio is a boolean on the video generation request; omitted, it defaults to the model's audio capability. Sume returns a finished MP4 as a hosted artifact. The docs describe no separate stems or per-track volume, so voice, ambience and music arrive mixed. If you need to rebalance, detach the audio with the audio detach tool and work on it in a timeline, which is a separate step.

Seedance 2.x also accepts audio references: the docs say audio and video references are honored by the Seedance 2.x models, and the Video 1.0 page says reference audio needs at least one image or video reference alongside it. A reference sample is for steering, not a guarantee of an exact voice.

How do you iterate without overspending?

Change one layer at a time. Render at 480p, 4 to 6 seconds, keep the picture sentence identical, and alter only the ambience line, then only the music line. Listening to three short drafts costs less than one 1080p failure, and Sume holds the estimated cost at submit and refunds it if a job fails.

When the sound is right, rerun the winning prompt at the final resolution. If generated speech keeps missing, set generate_audio to false and add narration, or use a clean render and burn authored captions with the video captions API, which accepts your own cues and skips speech-to-text.

Sources

Related posts

More in Models

All Models posts

Written by Sume