Make an AI voiceover louder: generation_config volume and a peak check
Sume TTS takes generation_config.volume from 0.5 to 2. Raise it, then check the WAV peak in Python so you never ship a clipped voiceover.

To make a Sume text-to-speech voiceover louder, set generation_config.volume on the request. The range is 0.5 to 2, with 1 as the neutral value. Louder is not always better: a voice pushed to 2 can clip, which sounds like crackle on the loud syllables. So render the take as a WAV, read its peak, and only then pick the setting. The script below renders the same line at 1.0 and 1.5 and prints the peak of each in dBFS, using nothing but the Python standard library.
Voice models are being sold on how they sound and how fast they start; Microsoft's MAI-Voice-2.1 and Flash launched on 2026-10-01 at $22 and $15 per 1M characters (October tracker, read 2026-10-06). Level is the plain, boring part of the pipeline that decides whether a voiceover is usable in a mix, so it is worth five minutes.
Render two takes and measure
import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
import array, math, urllib.request as ur, wave, io
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
def render(volume):
job = call("https://api.sume.com/v1/tts-router/generate", {
"model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
"transcript": "Order today and get free delivery on every pack.",
"generation_config": {"volume": volume},
"output_format": {"container": "wav", "encoding": "pcm_s16le"}}, f"level-demo-v1-{volume}")
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
arts = call(job["result_url"])["result"]["artifacts"]
return next(a["url"] for a in arts if a["type"] == "audio")
for volume in (1.0, 1.5):
url = render(volume)
with wave.open(io.BytesIO(ur.urlopen(url).read())) as w:
samples = array.array("h", w.readframes(w.getnframes()))
peak = max(abs(s) for s in samples)
print(volume, url, f"peak {20 * math.log10(peak / 32768):.1f} dBFS", "CLIPPED" if peak >= 32767 else "ok")How to read the result
A peak at or near 0 dBFS means the loudest sample hit the ceiling. The script flags a peak of 32767, the largest 16-bit value. A peak of -3 dBFS leaves headroom for a music bed or a limiter downstream. Pick the highest volume that stays clear of the ceiling by a couple of decibels, and keep that number in your voice profile so every episode gets the same level.
Two things the script does not do. It does not measure loudness, only peak: two files with the same peak can sound very different, and Sume's docs make no loudness-normalization claim for TTS output. If a platform asks for a loudness target, measure the final mix with a loudness meter. And it uses pcm_s16le, set explicitly, so the integer reading is correct; with the default MP3 you would be measuring decoded audio instead.
Cost and reruns
Each take is a separate job and a separate charge. The line above is 49 characters, so each take costs a fraction of a cent at $0.0475 per 1,000 characters (pricing, read 2026-10-06). The idempotency key includes the volume, so rerunning the script returns the finished takes and does not pay again. If you change the volume and keep the key, the API treats it as a conflict, not a new job. Poll with the fields in Jobs and results.
| generation_config.volume | What it is | When to try it |
|---|---|---|
| 0.5 | The quietest allowed value | Voice under loud music |
| 1.0 | Neutral | The default to compare against |
| 1.5 | Louder | When the take sounds thin next to a music bed |
| 2.0 | The loudest allowed value | Check the peak first; clipping is likely on loud voices |
Record the settings
Keep the take you choose and record its settings next to the script: model, voice, speed and volume. A voiceover that is re-rendered six weeks later with a different volume will not match the episodes around it, and the mismatch is easy to hear when two clips play back to back. Store the request body, not just the audio, so a re-render is a copy and not a guess.
The other levers
Use the same fields for other levers. speed runs from 0.6 to 1.5, and emotion goes in the same generation_config object. Sume has no SSML field, so there is no per-word volume; if one sentence is quieter than the rest, make it its own job with its own volume. See Sume TTS has no SSML field for what to use instead.
Sources
Related posts
More in Developers
- MATLAB webwrite and weboptions: generate an AI video and save the MP4
A MATLAB script posts a Wan 3.0 job to Sume with webwrite, polls the job status with webread and saves the finished clip with websave.
- Music-only Short: Timeline silence mode with a soundtrack bed
Build a Short with no voice: audio mode silence plus a soundtrack. The bed plays alone at the default -16 dB with no amix loss. Fields, limits, tail rule.
- Nearest supported aspect ratio per Sume image model, in Python
A model rejects an aspect ratio it does not list. Read the catalog, pick the closest ratio it accepts, generate, then crop to the exact shape you need.
- NestJS raw body controller to verify an AI video webhook signature
Enable rawBody in NestFactory, read req.rawBody in a controller and check Sume's sume-v1 HMAC before you act on a job.completed video event.
Written by Sume