Make an AI voiceover louder: generation_config volume and a peak check

Sume TTS takes generation_config.volume from 0.5 to 2. Raise it, then check the WAV peak in Python so you never ship a clipped voiceover.

5 min readSume
All posts

To make a Sume text-to-speech voiceover louder, set generation_config.volume on the request. The range is 0.5 to 2, with 1 as the neutral value. Louder is not always better: a voice pushed to 2 can clip, which sounds like crackle on the loud syllables. So render the take as a WAV, read its peak, and only then pick the setting. The script below renders the same line at 1.0 and 1.5 and prints the peak of each in dBFS, using nothing but the Python standard library.

Voice models are being sold on how they sound and how fast they start; Microsoft's MAI-Voice-2.1 and Flash launched on 2026-10-01 at $22 and $15 per 1M characters (October tracker, read 2026-10-06). Level is the plain, boring part of the pipeline that decides whether a voiceover is usable in a mix, so it is worth five minutes.

Render two takes and measure

import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
import array, math, urllib.request as ur, wave, io
def call(url, body=None, key=None):
    h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
    if key: h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body else None
    return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]

def render(volume):
    job = call("https://api.sume.com/v1/tts-router/generate", {
        "model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
        "transcript": "Order today and get free delivery on every pack.",
        "generation_config": {"volume": volume},
        "output_format": {"container": "wav", "encoding": "pcm_s16le"}}, f"level-demo-v1-{volume}")
    while not job["terminal"]:
        time.sleep(job.get("next_poll_after_seconds") or 2)
        job = call(job["status_url"])
    arts = call(job["result_url"])["result"]["artifacts"]
    return next(a["url"] for a in arts if a["type"] == "audio")

for volume in (1.0, 1.5):
    url = render(volume)
    with wave.open(io.BytesIO(ur.urlopen(url).read())) as w:
        samples = array.array("h", w.readframes(w.getnframes()))
    peak = max(abs(s) for s in samples)
    print(volume, url, f"peak {20 * math.log10(peak / 32768):.1f} dBFS", "CLIPPED" if peak >= 32767 else "ok")

How to read the result

A peak at or near 0 dBFS means the loudest sample hit the ceiling. The script flags a peak of 32767, the largest 16-bit value. A peak of -3 dBFS leaves headroom for a music bed or a limiter downstream. Pick the highest volume that stays clear of the ceiling by a couple of decibels, and keep that number in your voice profile so every episode gets the same level.

Two things the script does not do. It does not measure loudness, only peak: two files with the same peak can sound very different, and Sume's docs make no loudness-normalization claim for TTS output. If a platform asks for a loudness target, measure the final mix with a loudness meter. And it uses pcm_s16le, set explicitly, so the integer reading is correct; with the default MP3 you would be measuring decoded audio instead.

Cost and reruns

Each take is a separate job and a separate charge. The line above is 49 characters, so each take costs a fraction of a cent at $0.0475 per 1,000 characters (pricing, read 2026-10-06). The idempotency key includes the volume, so rerunning the script returns the finished takes and does not pay again. If you change the volume and keep the key, the API treats it as a conflict, not a new job. Poll with the fields in Jobs and results.

Volume settings for a first listening test (read 2026-10-06)
generation_config.volumeWhat it isWhen to try it
0.5The quietest allowed valueVoice under loud music
1.0NeutralThe default to compare against
1.5LouderWhen the take sounds thin next to a music bed
2.0The loudest allowed valueCheck the peak first; clipping is likely on loud voices

Record the settings

Keep the take you choose and record its settings next to the script: model, voice, speed and volume. A voiceover that is re-rendered six weeks later with a different volume will not match the episodes around it, and the mismatch is easy to hear when two clips play back to back. Store the request body, not just the audio, so a re-render is a copy and not a guess.

The other levers

Use the same fields for other levers. speed runs from 0.6 to 1.5, and emotion goes in the same generation_config object. Sume has no SSML field, so there is no per-word volume; if one sentence is quieter than the rest, make it its own job with its own volume. See Sume TTS has no SSML field for what to use instead.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume