Same volume every Shorts episode: gain_db, duck_db, one music bed

Sume does not measure loudness. To keep a series consistent, pin audio.gain_db, soundtrack gain_db and duck_db in one function with one bed. Python plan loop.

6 min readSume
All posts

What Sume can and cannot do for series loudness

If every episode of a Shorts series should sound as loud as the last, the documented tools are fixed numbers, not measurement. Timeline 1.0 lets you set audio.gain_db on the spine from -60 to 12, soundtrack.gain_db on the music bed, and soundtrack.duck_db from 0 to 20 to lower the bed under the spine. The docs describe no loudness analysis, LUFS target or normalization step, so we will not claim one. What you get is determinism: the same numbers in, the same mix rules out, for every episode.

That is worth doing now because a series has an audience that binges. The October platform roundup says YouTube's Shorts series, with seasons, episodes and sequential playback, began rolling out on 23 September. When episodes play back to back, a 6 dB jump between them is the first thing a viewer notices.

Put the mix in one function

The cheapest consistency tool is a single function that builds every episode body. Keep the music bed identical across the season, set its gain_db low enough that the spine leads, and let duck_db lower it further while someone speaks. Ducking needs a real spine, so it works for voiceover episodes and is refused for silent ones with duck_requires_audio_spine. Set loop: true so a short bed covers a long episode and fade_out_seconds (at most 10) so it ends cleanly.

Pin one more thing: render all episodes from spines made the same way. If episode 3's voice came from a different recording chain than episode 2's, no constant will fix it. For voiceover made with Sume's own speech jobs, reuse the same voice and model for the whole season so the level you set once keeps meaning the same thing.

Values to pin and where they live

These fields and ranges come from the Timeline 1.0 program table.

Audio fields for a consistent series mix, from the Timeline 1.0 docs (read 2026-10-06)
FieldRange or ruleRole in the mix
audio.gain_db-60 to 12, not with silence modeLevel of the spine, the voice
soundtrack.gain_dbSet per requestLevel of the music bed
soundtrack.duck_db0 to 20, needs a real spineHow far the bed drops under the spine
soundtrack.looptrue or falseRepeats a short bed under a long episode
soundtrack.fade_out_secondsUp to 10 sEnds the bed without a hard stop
output.fade_in_seconds, fade_out_seconds0 to 5 s each, sum within output lengthFades of the whole render

Python: plan three episodes with the same mix

This plans three episodes of different lengths through one body builder, so only the voice file and the length vary. The plan is unbilled, needs no idempotency key and returns billable_minutes for each. Replace the example URLs with your own imported artifacts, and POST the same bodies to /v1/timeline-1.0/render with unique keys when you are happy.

import json, os, time, urllib.request

def sume(method, path, body=None, key=None):
    h = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
         "Content-Type": "application/json", "User-Agent": "sume-example/1.0"}
    if key:
        h["Idempotency-Key"] = key
    data = json.dumps(body).encode() if body is not None else None
    req = urllib.request.Request("https://api.sume.com" + path, data, h, method=method)
    with urllib.request.urlopen(req) as r:
        return json.load(r)

def episode_body(voice_url, bed_url, seconds):
    return {"audio": {"url": voice_url, "duration_seconds": seconds, "gain_db": 0},
            "video": [{"source_url": "https://media.sume.com/artifacts/artf_demo/bg.mp4",
                       "start": 0, "duration": seconds}],
            "soundtrack": {"url": bed_url, "gain_db": -18, "duck_db": 12,
                           "loop": True, "fade_out_seconds": 2}}

bed = "https://media.sume.com/artifacts/artf_demo/theme.mp3"
for n, secs in ((1, 41), (2, 52), (3, 38)):
    body = episode_body(f"https://media.sume.com/artifacts/artf_demo/ep{n}.wav", bed, secs)
    print(n, sume("POST", "/v1/timeline-1.0/plan", body)["billable_minutes"])

Choosing the numbers

We deliberately do not give you a magic pair of values, because the right gain depends on your voice recording and your bed. Choose by ear on one episode, write the numbers into the function, and treat them as the season's mix spec. A workable routine is to render episode 1, listen on a phone speaker, adjust soundtrack.gain_db and duck_db until the voice leads without the bed disappearing, then freeze the values. Every later episode inherits them, so the only variable left is the voice file.

When a later episode sounds different, the fix is to find what changed upstream, not to nudge the numbers per episode. Per-episode nudges are how a season ends up with nine different mixes. If you must change the spec, change it in the function, re-plan the whole season, and decide whether to re-render the older episodes, using the plan to cost that decision first.

What the render returns

A finished timeline job returns a new MP4 on media.sume.com with duration_seconds, segment_count and billable_minutes, plus any soft warnings such as a padded short source. Under the documented $0.10 per whole output minute, three episodes under 60 seconds bill three minutes, or $0.30 in total; confirm the live rate in GET /v1/catalog. Keep the plan output next to each render so the cost you quoted and the cost you paid can be compared line by line.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume