AI voiceover too loud: three Sume gain knobs and what each one costs

TTS generation_config.volume (0.5-2), Timeline audio.gain_db (-60 to 12) and soundtrack.duck_db (0-20). Which to change, and which means paying for new audio.

5 min readSume
All posts

If an AI voiceover is too loud against music in a Sume render, change soundtrack.gain_db or soundtrack.duck_db in the Timeline request first. They cost nothing extra because you are rendering anyway. Change the voice itself with audio.gain_db (-60 to 12 dB) in the same render. Only generation_config.volume (0.5 to 2) on the TTS request means a new audio file, a new job and new characters at $0.0475 per 1,000.

The three places

Sume has loudness controls at two stages. At synthesis, the TTS request accepts generation_config.volume from 0.5 to 2, and the finished job records it back, null if you did not send it. At assembly, the Timeline render accepts audio.gain_db for the spine and, for the optional bed, soundtrack.gain_db and duck_db from 0 to 20, which lowers the bed under speech and needs a real spine, not silence.

Loudness controls and what changing each one costs, rates read 2026-10-09
ControlStageRangeCost to change
generation_config.volumeTTS request0.5 to 2New TTS job: $0.0475 per 1,000 characters
audio.gain_dbTimeline render-60 to 12 dB, not with silenceRe-render: $0.10 per ceil(output minute)
soundtrack.gain_dbTimeline rendera dB valueRe-render
soundtrack.duck_dbTimeline render0 to 20, needs a real spineRe-render

A worked example

A 45-second short has a 600-character voiceover and a bed. The voice cost 600 x $0.0475 / 1,000 = $0.0285. The render reserves 1 minute at $0.10. If the mix is wrong, a render-side fix costs $0.10. A new voice take with a different volume costs $0.0285 plus the same $0.10 re-render, $0.1285, and the take may not sound identical. So fix loudness in the mix, and spend a new take only when the delivery itself is wrong.

Run the unbilled plan call first when you only need to confirm length: it returns billable_minutes and estimated_cost_usd_micros and creates no job.

Why the spine stays untouched

The voiceover that you render is the audio spine, and it is a file, not a setting. That is why the cheap fixes live in the render: the spine is read as is, and gain, bed level and ducking are applied on top at assembly. A voice that was synthesised at the right level and then mixed down for the bed leaves you the cleanest source if you later re-cut the video for another platform. Keep the original TTS file and the Timeline request side by side, so a different mix is a re-render, not a re-record.

Rules of thumb

Four rules keep a mix cheap to fix.

  • Leave the TTS volume unset unless the clip clips or whispers. Both gain stages downstream are cheaper to correct.
  • Duck first, then lower the bed. duck_db reacts to speech and leaves the bed full level between lines.
  • Do not push audio.gain_db to the top of its range to rescue a quiet take; a take that was quiet at synthesis may also be noisy.
  • Join files in wav. The docs say mp3 adds priming padding at each edge, which shows up as a faint gap when files are joined.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume