Fit narration to a fixed slot: measure first, then set TTS speed

A 45-second cap or a 30-second slot decides your script. Render once with word timings, compute the speed ratio, and rewrite only if outside 0.6 to 1.5.

6 min readSume
All posts

How do you make a voice-over exactly fit a fixed time slot? Measure it, do not guess from word count. Render the script once at normal speed with word timestamps, read the end of the last word, divide by the slot, and either set generation_config.speed or cut the script.

Slots are everywhere: Microsoft's announcement (read 2026-10-04) says MAI-Voice-2.1-Flash generates 45 seconds of audio per call, and Sume's avatar videos accept scripts that estimate to 4 to 60 seconds, per the avatar video guide.

The speed range

Sume's TTS contract takes generation_config.speed between 0.6 and 1.5, and volume between 0.5 and 2.0. The older speed enum (slow, normal, fast) is deprecated in favour of the numeric field. A multiplier of 1.25 means the take should run in about 80 percent of the time, but the contract does not promise exact scaling, so measure again after you change it.

Measure, then set

This script renders once, reads the timings, and prints the speed to try next. If it falls outside 0.6 to 1.5 it tells you to edit the script instead.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def wait(d):
    while not d["result_ready"]:
        if d.get("terminal"):
            raise RuntimeError(d.get("status"))
        time.sleep(d.get("next_poll_after_seconds") or 3)
        j = requests.get(d["status_url"], headers=H).json()
        d = j.get("data", j)
    j = requests.get(d["result_url"], headers=H).json()
    return j.get("data", j)["result"]

SLOT = 30.0
script = "Your script goes here."
body = {"transcript": script, "avatar_handle": "your_handle",
        "timestamps": {"words": True},
        "generation_config": {"speed": 1.0}}
r = requests.post(API + "/v1/tts-1.0/generate", json=body,
                  headers={**H, "Idempotency-Key": "fit-measure-001"})
r.raise_for_status()
result = wait(r.json()["data"])
spoken = result["words"][-1]["end"]
speed = round(spoken / SLOT, 2)
print(spoken, speed, "ok" if 0.6 <= speed <= 1.5 else "edit the script")

Rewrite or speed up

Past about 1.25 most voices start to sound hurried, which is a listening judgement, not a contract limit. If the ratio is above that, cut words. If it is far below 1.0, you are padding: add a sentence or leave the picture on screen longer.

Keep the measuring render. The job records the voice, language and settings it was made with, per Jobs and results, so the final take can reuse them and differ only in speed.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume