Match loudness between voiceover takes: test sentence plus volume
Two takes recorded on different days sound different in level. Measure one fixed test sentence, then set Sume TTS generation_config.volume (0.5 to 2.0).

To match loudness between two voiceover takes, render the same fixed test sentence on each day, measure its mean level, and move generation_config.volume by the gap. Sume documents volume as a multiplier from 0.5 to 2.0, and says nothing in the pages read about loudness normalising TTS output, so two takes can land at different levels. If the multiplier scales amplitude, 2.0 is about +6 dB and 0.5 about -6 dB; that is an assumption, so treat the first value as a guess and measure again. A gap larger than that range needs gain after generation.
The volume range comes from the TTS schema in the Sume API reference; the post-generation gain range from Timeline 1.0; the measuring filter from FFmpeg's volumedetect documentation, read on 2026-10-03. Streaming-platform loudness targets are in podcast loudness; this page is about two takes drifting.
Why do two takes of the same voice differ in level?
Each take is a separate generation of a different script. Sume's docs promise the settings you send, not an identical level, so a quiet Monday take and a louder Friday take can both be correct output. Judge them on the same words instead of on whole scripts, which differ in length and content.
How do I measure and set the volume?
Pin one sentence, such as "Welcome back. Here is this week's update.", and generate it at volume: 1.0 with the same voice, speed and wav format every time. Keep the first one as the reference. FFmpeg's volumedetect filter reports mean volume (RMS) and maximum volume in decibels relative to full scale. This script reads both files and suggests a multiplier for the new take, clamped to Sume's range.
import re, subprocess, sys
def mean_db(path):
log = subprocess.run(["ffmpeg", "-hide_banner", "-i", path, "-af", "volumedetect",
"-f", "null", "-"], capture_output=True, text=True).stderr
return float(re.search(r"mean_volume: (-?[\d.]+) dB", log).group(1))
ref, new, vol = mean_db(sys.argv[1]), mean_db(sys.argv[2]), float(sys.argv[3])
gap = ref - new
want = vol * 10 ** (gap / 20)
print(f"reference {ref} dB, new {new} dB, gap {gap:+.1f} dB")
print("try volume", round(min(2.0, max(0.5, want)), 2))
if abs(gap) > 6:
print("gap exceeds about 6 dB: fix it after generation with Timeline gain_db")What are the limits?
Run the script as python match.py reference.wav new.wav 1.0, regenerate the test sentence at the suggested value, and measure once more.
- Raising the level can clip, so check that
max_volumestays below 0 dB. - Keep one
volumeper project once matched, and re-test only when the voice or format changes.
| Control | Where | Range |
|---|---|---|
| generation_config.volume | TTS request | 0.5 to 2.0 multiplier |
| audio.gain_db | Timeline render spine | -60 to 12 dB |
| soundtrack.duck_db | Timeline render music bed | 0 to 20 dB |
Sources
Related posts
More in Developers
- MAX_MCP_OUTPUT_TOKENS 25,000: size a Sume jobs_wait wave read
Claude Code caps MCP output at 25,000 tokens and saves larger results to a file. How that meets Sume's jobs_wait include_results and a batch jobs_result read.
- MCP 401 challenge: the WWW-Authenticate header Sume returns
What a client sees when it calls Sume's MCP endpoint with no token: the WWW-Authenticate challenge, its resource_metadata URL and scope, and the spec.
- mcp_health on Sume: which credential and scopes is this session using?
Sume's mcp_health tool reports the auth source and the credential behind a session: an OAuth token with client and scopes, or an API key prefix. How to read it.
- Same Sume MCP URL, different tools: what your credential can see
Two clients on one Sume MCP URL can list different tools. Credential scope sets the surface, so compare mcp_health tools[] before blaming the client.
Written by Sume