Fit narration to a fixed slot: measure first, then set TTS speed
A 45-second cap or a 30-second slot decides your script. Render once with word timings, compute the speed ratio, and rewrite only if outside 0.6 to 1.5.

How do you make a voice-over exactly fit a fixed time slot? Measure it, do not guess from word count. Render the script once at normal speed with word timestamps, read the end of the last word, divide by the slot, and either set generation_config.speed or cut the script.
Slots are everywhere: Microsoft's announcement (read 2026-10-04) says MAI-Voice-2.1-Flash generates 45 seconds of audio per call, and Sume's avatar videos accept scripts that estimate to 4 to 60 seconds, per the avatar video guide.
The speed range
Sume's TTS contract takes generation_config.speed between 0.6 and 1.5, and volume between 0.5 and 2.0. The older speed enum (slow, normal, fast) is deprecated in favour of the numeric field. A multiplier of 1.25 means the take should run in about 80 percent of the time, but the contract does not promise exact scaling, so measure again after you change it.
Measure, then set
This script renders once, reads the timings, and prints the speed to try next. If it falls outside 0.6 to 1.5 it tells you to edit the script instead.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def wait(d):
while not d["result_ready"]:
if d.get("terminal"):
raise RuntimeError(d.get("status"))
time.sleep(d.get("next_poll_after_seconds") or 3)
j = requests.get(d["status_url"], headers=H).json()
d = j.get("data", j)
j = requests.get(d["result_url"], headers=H).json()
return j.get("data", j)["result"]
SLOT = 30.0
script = "Your script goes here."
body = {"transcript": script, "avatar_handle": "your_handle",
"timestamps": {"words": True},
"generation_config": {"speed": 1.0}}
r = requests.post(API + "/v1/tts-1.0/generate", json=body,
headers={**H, "Idempotency-Key": "fit-measure-001"})
r.raise_for_status()
result = wait(r.json()["data"])
spoken = result["words"][-1]["end"]
speed = round(spoken / SLOT, 2)
print(spoken, speed, "ok" if 0.6 <= speed <= 1.5 else "edit the script")
Rewrite or speed up
Past about 1.25 most voices start to sound hurried, which is a listening judgement, not a contract limit. If the ratio is above that, cut words. If it is far below 1.0, you are padding: add a sentence or leave the picture on screen longer.
Keep the measuring render. The job records the voice, language and settings it was made with, per Jobs and results, so the final take can reuse them and differ only in speed.
Sources
Related posts
More in Developers
- format_not_forkable 409 in the Sume API: what it means and the fix
Sume returns 409 format_not_forkable when the id you called names a built-in capability, not a Format card. How to tell, and which ids to call instead.
- Gemini 3.8 Live audio: wrap 24 kHz PCM in WAV, resample to 16 kHz
Gemini 3.8 Live takes 16-bit 16 kHz PCM in and returns 24 kHz out. A Python WAV wrapper, an ffmpeg resample command, and the Sume detach settings that match.
- Gemini video understanding 88% fewer tokens vs Sume Video inspect
Gemini reports up to 88% fewer tokens on long video. Sume Video inspect and Reference ingest take another route: stills, transcript and a manifest.
- Gemini CLI 0.62 MCP titles: reading Sume's tool names
Gemini CLI v0.62.0 formats MCP tool call titles as structured signatures. Sume tool ids are underscore names such as generate_image; dotted aliases map to them.
Written by Sume