Does generate_audio change Seedance 2.0's price?

No. fal says Seedance 2.0 audio costs the same on or off, and Sume's Seedance price has no audio input: seedance-2 reserves about $0.38 per 720p 16:9 second.

4 min readSume
All posts

No. fal's Seedance 2.0 pages say audio costs the same whether generate_audio is on or off, and Sume's price estimate for the Seedance ids has no audio input at all: the reserve comes from resolution, aspect ratio, duration and whether a reference video is attached. seedance-2 reserves about $0.38 per 720p 16:9 second either way, before the agent fee.

The vendor facts were read 2026-09-29. Sume's side comes from the Video generation docs and the pricing code.

What does fal say about audio and price?

Two fal pages state it in different words. Both are about the Seedance 2.0 endpoints, not about Sume.

From fal's fast text-to-video and reference-to-video pages, and Sume's Video generation docs, read 2026-09-29.
SourceWhat it says about audioPrice effect
fal, fast text-to-videogenerate_audio defaults to true: synchronized SFX, ambient sound and lip-synced speech."Same price either way."
fal, reference-to-videoAudio generation is included regardless of the setting.No extra cost
Sume, generate_audioOptional boolean. Defaults to the model's audio capability.Not an input to the Seedance estimate; seedance-2 is about $0.38 per 720p 16:9 second

What does generate_audio do on Sume?

The docs describe generate_audio as whether to generate audio alongside the video, defaulting to the model's audio capability. The catalog entry for seedance-2 shows generate_audio: true, so leaving the field out gives the model's default. The API refuses generate_audio: true only for a model whose descriptor says it cannot make audio; it refuses false for MiniMax H3, MiniMax H3 Max and Gemini Omni Flash 1.1, which always make audio. Seedance is in neither group in that code.

Turning audio off therefore does not lower the reserve. Turn it off when you plan to lay your own soundtrack over the clip, not to change the bill.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "seedance-2",
    "prompt": "A barista pours latte art, close-up, quiet cafe",
    "duration": 6,
    "resolution": "720p",
    "aspect_ratio": "16:9",
    "generate_audio": false
  }'

How do I ask for spoken dialogue?

fal's Seedance 2.0 text-to-video page says to put spoken dialogue in double quotes for lip-synced audio, and that the model then generates matching lip movements and voice. That is fal's prompt guidance for the endpoint it hosts. Sume's docs do not describe a prompt syntax, so try it with a short line first: A chef looks at the camera and says "The sauce needs five more minutes." Because audio does not change the reserve, a short test with audio costs the same as the same clip with audio off.

What can change the price of a Seedance 2.0 clip?

Resolution, aspect ratio and duration set the frame size and the token count. The model id sets the rate per 1,000 tokens. A reference video adds input duration to the billable tokens. Read the amount actually charged from the job's usage.cost; the docs describe it as the Sume billable amount.

Sources

Related posts

More in Models

All Models posts

Written by Sume