Sume TTS volume 0.5 to 2.0: set the voiceover level before the mix

generation_config volume is a multiplier from 0.5 to 2.0 on Sume TTS. Use it to match narration loudness across jobs before you join or mix them.

5 min readSume
All posts

Sume TTS lets you set loudness per request with generation_config.volume, a multiplier from 0.5 to 2.0. Use it when several jobs will be joined into one track and one of them comes out noticeably quieter or louder than the rest. A value of 1 is the neutral multiplier, 0.5 halves the level and 2.0 doubles it. Keep changes small and listen to the result.

Set it on the request, then check the joined track, rather than fixing levels after the fact.

What does the spec define?

In the Sume OpenAPI spec, volume is a number with a minimum of 0.5 and a maximum of 2, described as a volume multiplier in [0.5, 2.0]. It lives beside speed (0.6 to 1.5) and emotion in generation_config, which accepts no other keys.

Volume multiplier values, Sume spec read 2026-10-04
volumeMeaning
0.5Half the level, the minimum
1.0Neutral multiplier
1.5One and a half times the level
2.0Double, the maximum
Below 0.5 or above 2.0Rejected by validation

When should you change it?

Mostly when you are matching jobs, not making a single clip louder. Split scripts produce separate files, and a mismatch is easy to hear at the join.

  • Render a short test sentence from each voice and compare.
  • Change only one job's volume and keep the rest at the default.
  • For music under speech, adjust the music in the timeline instead; see the ducking example linked below.

How do you set it?

Add volume to the request body:

import os, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
volume = float(os.environ.get("VOICE_VOLUME", "1.2"))
assert 0.5 <= volume <= 2.0, "volume must be within 0.5 to 2.0"

r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
    headers=H, timeout=60,
    json={
        "transcript": "Part two of the walkthrough.",
        "voice": {"id": os.environ["VOICE_ID"]},
        "language": "en",
        "generation_config": {"volume": volume},
    })
r.raise_for_status()
print(r.json()["data"]["job"]["id"])

Where does the mix happen?

Join the parts with a timeline audio concat, and handle music levels as shown in setting duck dB in a Sume timeline.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume