Fit narration to a 30-second slot: tune TTS speed from measured length

Measure a TTS take with timestamps.words, compute the speed that hits 30 seconds, and know when to cut words instead. Python, within Sume's 0.6 to 1.5 range.

4 min readSume
All posts

Request word timings with the take, read the end time of the last word, and divide: new speed is the old speed times measured seconds over target seconds. A 34.2 second take at speed 1.0 needs about 1.14 for a 30 second slot. Sume TTS speed runs from 0.6 to 1.5, so a take that needs more than that must be rewritten, not sped up.

The formula assumes length scales with 1 / speed. That is our assumption, so confirm the retake's real length.

How do I measure the take?

Send timestamps: { words: true } with the TTS request. The completed result then includes words[] with monotonic start and end seconds. The last word's end is your speech length, which is a better number than the file duration if the file carries silence at the edges.

What speed do I try next?

The function below clamps to the documented range and flags a take that cannot be fixed by speed alone. It runs as written.

LOW, HIGH = 0.6, 1.5   # generation_config.speed range

def next_speed(speed, measured, target):
    wanted = speed * measured / target      # assumes length scales with 1/speed
    return min(HIGH, max(LOW, round(wanted, 2))), wanted

for speed, measured in [(1.0, 34.2), (1.0, 50.0), (1.0, 20.0)]:
    clamped, wanted = next_speed(speed, measured, 30.0)
    note = "" if clamped == round(wanted, 2) else "  <- clamped; edit the script instead"
    print(f"{measured}s at {speed} -> try {clamped} (wanted {wanted:.2f}){note}")

When should I cut words instead?

Speech sped to the top of the range sounds rushed, and the clamp case above shows a 50 second take that would need 1.67. Cut the script to roughly the target length instead. The words-per-minute post in the related list gives a quick way to estimate that before spending a take.

Between those extremes, change one thing per retake. Each completed take is a new generation of the same characters, so budget two or three tries per script and stop when the result is within a half second of the slot.

Worked examples for a 30 s target (our arithmetic, read 2026-10-02).
Measured at speed 1.0Wanted speedAction
34.2 s1.14Retake at 1.14
30.4 s1.01Keep, trim 0.4 s in Timeline
50.0 s1.67Out of range: cut the script
20.0 s0.67Retake at 0.67 or add a beat

Does the slot have to match exactly?

Timeline video slots may trail the audio spine by at most 0.5 seconds, so a take within about half a second of the slot can be fitted in the render. Anything longer than that, change the take.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume