Match loudness between voiceover takes: test sentence plus volume

Two takes recorded on different days sound different in level. Measure one fixed test sentence, then set Sume TTS generation_config.volume (0.5 to 2.0).

4 min readSume
All posts

To match loudness between two voiceover takes, render the same fixed test sentence on each day, measure its mean level, and move generation_config.volume by the gap. Sume documents volume as a multiplier from 0.5 to 2.0, and says nothing in the pages read about loudness normalising TTS output, so two takes can land at different levels. If the multiplier scales amplitude, 2.0 is about +6 dB and 0.5 about -6 dB; that is an assumption, so treat the first value as a guess and measure again. A gap larger than that range needs gain after generation.

The volume range comes from the TTS schema in the Sume API reference; the post-generation gain range from Timeline 1.0; the measuring filter from FFmpeg's volumedetect documentation, read on 2026-10-03. Streaming-platform loudness targets are in podcast loudness; this page is about two takes drifting.

Why do two takes of the same voice differ in level?

Each take is a separate generation of a different script. Sume's docs promise the settings you send, not an identical level, so a quiet Monday take and a louder Friday take can both be correct output. Judge them on the same words instead of on whole scripts, which differ in length and content.

How do I measure and set the volume?

Pin one sentence, such as "Welcome back. Here is this week's update.", and generate it at volume: 1.0 with the same voice, speed and wav format every time. Keep the first one as the reference. FFmpeg's volumedetect filter reports mean volume (RMS) and maximum volume in decibels relative to full scale. This script reads both files and suggests a multiplier for the new take, clamped to Sume's range.

import re, subprocess, sys
def mean_db(path):
    log = subprocess.run(["ffmpeg", "-hide_banner", "-i", path, "-af", "volumedetect",
                          "-f", "null", "-"], capture_output=True, text=True).stderr
    return float(re.search(r"mean_volume: (-?[\d.]+) dB", log).group(1))
ref, new, vol = mean_db(sys.argv[1]), mean_db(sys.argv[2]), float(sys.argv[3])
gap = ref - new
want = vol * 10 ** (gap / 20)
print(f"reference {ref} dB, new {new} dB, gap {gap:+.1f} dB")
print("try volume", round(min(2.0, max(0.5, want)), 2))
if abs(gap) > 6:
    print("gap exceeds about 6 dB: fix it after generation with Timeline gain_db")

What are the limits?

Run the script as python match.py reference.wav new.wav 1.0, regenerate the test sentence at the suggested value, and measure once more.

  • Raising the level can clip, so check that max_volume stays below 0 dB.
  • Keep one volume per project once matched, and re-test only when the voice or format changes.
Level controls and their ranges, from Sume's API schema and Timeline docs, read 2026-10-03.
ControlWhereRange
generation_config.volumeTTS request0.5 to 2.0 multiplier
audio.gain_dbTimeline render spine-60 to 12 dB
soundtrack.duck_dbTimeline render music bed0 to 20 dB

Sources

Related posts

More in Developers

All Developers posts

Written by Sume