Sume TTS volume 0.5 to 2.0: set the voiceover level before the mix
generation_config volume is a multiplier from 0.5 to 2.0 on Sume TTS. Use it to match narration loudness across jobs before you join or mix them.

Sume TTS lets you set loudness per request with generation_config.volume, a multiplier from 0.5 to 2.0. Use it when several jobs will be joined into one track and one of them comes out noticeably quieter or louder than the rest. A value of 1 is the neutral multiplier, 0.5 halves the level and 2.0 doubles it. Keep changes small and listen to the result.
Set it on the request, then check the joined track, rather than fixing levels after the fact.
What does the spec define?
In the Sume OpenAPI spec, volume is a number with a minimum of 0.5 and a maximum of 2, described as a volume multiplier in [0.5, 2.0]. It lives beside speed (0.6 to 1.5) and emotion in generation_config, which accepts no other keys.
| volume | Meaning |
|---|---|
| 0.5 | Half the level, the minimum |
| 1.0 | Neutral multiplier |
| 1.5 | One and a half times the level |
| 2.0 | Double, the maximum |
| Below 0.5 or above 2.0 | Rejected by validation |
When should you change it?
Mostly when you are matching jobs, not making a single clip louder. Split scripts produce separate files, and a mismatch is easy to hear at the join.
- Render a short test sentence from each voice and compare.
- Change only one job's volume and keep the rest at the default.
- For music under speech, adjust the music in the timeline instead; see the ducking example linked below.
How do you set it?
Add volume to the request body:
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
volume = float(os.environ.get("VOICE_VOLUME", "1.2"))
assert 0.5 <= volume <= 2.0, "volume must be within 0.5 to 2.0"
r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
headers=H, timeout=60,
json={
"transcript": "Part two of the walkthrough.",
"voice": {"id": os.environ["VOICE_ID"]},
"language": "en",
"generation_config": {"volume": volume},
})
r.raise_for_status()
print(r.json()["data"]["job"]["id"])
Where does the mix happen?
Join the parts with a timeline audio concat, and handle music levels as shown in setting duck dB in a Sume timeline.
Sources
Related posts
More in Developers
- Check for a newer Sume CLI release in CI without auto-upgrading
sume update --check reports whether a newer GitHub Release exists and changes no files. Run it on a schedule, log it, and bump your pinned tag by pull request.
- Sume uploadFile with raw bytes needs a content type, or no request
uploadFile refuses a Uint8Array or ArrayBuffer without contentType and throws SumeUploadError at step create before any call. Pass a typed Blob or a MIME.
- Patching Supabase Postgres 17.11 vs Sume's 10-attempt webhook budget
Supabase's September 25 Postgres 15.19 and 17.11 releases fix 44 CVEs. A restart can outlast Sume's ten 30-second webhook attempts, so plan a redeliver.
- Supabase cached egress is $0.03/GB: cost of serving a 20 MB AI clip
Supabase lists cached Storage egress at $0.03 per GB. Worked arithmetic for serving generated clips, and when to link a Sume media URL instead of copying.
Written by Sume