AI voiceover too loud: three Sume gain knobs and what each one costs
TTS generation_config.volume (0.5-2), Timeline audio.gain_db (-60 to 12) and soundtrack.duck_db (0-20). Which to change, and which means paying for new audio.

If an AI voiceover is too loud against music in a Sume render, change soundtrack.gain_db or soundtrack.duck_db in the Timeline request first. They cost nothing extra because you are rendering anyway. Change the voice itself with audio.gain_db (-60 to 12 dB) in the same render. Only generation_config.volume (0.5 to 2) on the TTS request means a new audio file, a new job and new characters at $0.0475 per 1,000.
The three places
Sume has loudness controls at two stages. At synthesis, the TTS request accepts generation_config.volume from 0.5 to 2, and the finished job records it back, null if you did not send it. At assembly, the Timeline render accepts audio.gain_db for the spine and, for the optional bed, soundtrack.gain_db and duck_db from 0 to 20, which lowers the bed under speech and needs a real spine, not silence.
| Control | Stage | Range | Cost to change |
|---|---|---|---|
| generation_config.volume | TTS request | 0.5 to 2 | New TTS job: $0.0475 per 1,000 characters |
| audio.gain_db | Timeline render | -60 to 12 dB, not with silence | Re-render: $0.10 per ceil(output minute) |
| soundtrack.gain_db | Timeline render | a dB value | Re-render |
| soundtrack.duck_db | Timeline render | 0 to 20, needs a real spine | Re-render |
A worked example
A 45-second short has a 600-character voiceover and a bed. The voice cost 600 x $0.0475 / 1,000 = $0.0285. The render reserves 1 minute at $0.10. If the mix is wrong, a render-side fix costs $0.10. A new voice take with a different volume costs $0.0285 plus the same $0.10 re-render, $0.1285, and the take may not sound identical. So fix loudness in the mix, and spend a new take only when the delivery itself is wrong.
Run the unbilled plan call first when you only need to confirm length: it returns billable_minutes and estimated_cost_usd_micros and creates no job.
Why the spine stays untouched
The voiceover that you render is the audio spine, and it is a file, not a setting. That is why the cheap fixes live in the render: the spine is read as is, and gain, bed level and ducking are applied on top at assembly. A voice that was synthesised at the right level and then mixed down for the bed leaves you the cleanest source if you later re-cut the video for another platform. Keep the original TTS file and the Timeline request side by side, so a different mix is a re-render, not a re-record.
Rules of thumb
Four rules keep a mix cheap to fix.
- Leave the TTS
volumeunset unless the clip clips or whispers. Both gain stages downstream are cheaper to correct. - Duck first, then lower the bed.
duck_dbreacts to speech and leaves the bed full level between lines. - Do not push
audio.gain_dbto the top of its range to rescue a quiet take; a take that was quiet at synthesis may also be noisy. - Join files in wav. The docs say mp3 adds priming padding at each edge, which shows up as a faint gap when files are joined.
Sources
Related posts
More in Developers
- Alibaba Wan 3.0 Model Studio request to a Sume /v1/videos body
Map a Model Studio wan3.0-video call (input.media, parameters, X-DashScope-Async) to POST /v1/videos with model wan-3.0. Field by field.
- Wan 3.0 workspace-id region hosts vs Sume's one API base URL
Alibaba's Wan 3.0 URLs carry a workspace id and a region host, in six regions. Sume's video API is one base URL; what to change when you port the call.
- Failed AI video jobs: Google bills successes, Sume refunds the hold
Google says Veo videos are charged only when generated. Sume holds the price at submit, captures it on success and refunds it on failure or early cancel.
- Audio detach retry: what idempotency_hit true means for the $0.01
Retrying an audio detach with the same Idempotency-Key and body returns the first job (idempotency_hit true): $0.01 once. A changed body returns 409.
Written by Sume