Does Seedance 2.5 generate audio through the Sume API?

Yes: Seedance 2.5 on Sume has an optional generate_audio flag. ByteDance describes joint audio-video generation. What to send and how it is priced.

4 min readSume
All posts

Yes. Seedance 2.5 on Sume has an optional generate_audio boolean, and ByteDance describes the model as doing unified audio-video joint generation. If you omit the flag, the default follows the model's audio capability.

Sume prices Seedance on video tokens, so check usage.cost on a first job with audio on.

The flag

Send it to POST /v1/videos. Setting generate_audio to false asks for a silent clip.

{"model": "seedance-2.5",
 "prompt": "Street vendor frying dumplings, crowd chatter, sizzling oil",
 "duration": 12,
 "resolution": "720p",
 "aspect_ratio": "9:16",
 "generate_audio": true}

Audio as a reference

You can also provide audio references, up to 3 on Seedance 2.5. Use this to carry a voice or sound bed across clips.

Checking the catalog

Audio support in Sume's video docs (read 2026-10-07)
ModelAudio
seedance-2.5optional generate_audio; audio references
wan-3.0optional generate_audio; audio references
gemini-omni-flash-1.1always on; no audio references

What ByteDance says

The launch post describes Seedance 2.5 as unified multimodal audio-video joint generation: sound and picture come from one pass rather than a video model plus a separate sound step.

Good prompts for sound

Name the sound as part of the scene: ambience, a specific action noise, or a line of dialogue. If the clip has a speaker, say who and what tone. If you want only music, ask for no speech.

If you need precise, repeatable voice work, a dedicated text-to-speech step may suit better than relying on the video model for every line.

  • Add generate_audio: true for the first test.
  • Listen on headphones before judging.
  • Compare against a silent render for cost.

Sources

Related posts

More in Models

All Models posts

Written by Sume