Fit a voiceover to a 30-second ad: measure pace, then set TTS speed

Send timestamps.words, read the last word end time, and rescale generation_config.speed (0.6 to 1.5) until a Sume TTS take fits its slot. Two takes, 2 cents.

4 min readSume
All posts

An ad slot is 30.0 seconds, and the script reads 34. Trimming copy is the better fix, but sometimes the legal line stays and the speed has to move. Sume TTS 1.0 can measure the length before you commit to a final take: ask for timestamps: {"words": true} and the result lists every word with start and end. The end of the last word is the real spoken length, with no guesswork from a word count.

The method

Render the script once at the default speed. Read the last word's end. Divide it by your target to get the ratio. generation_config.speed runs from 0.6 to 1.5, so a ratio of 34 / 29.5 = 1.15 asks for speed 1.15. Leave half a second under the slot for the tail. Render a second take at that speed and check again. Two takes of a 450-character script quote 3 cents each at $0.0475 per 1,000 characters after rounding, so the loop costs 6 cents.

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

def length(text, speed=None):
    body = {"transcript": text, "language": "en", "timestamps": {"words": True},
            "voice": {"id": os.environ["VOICE_ID"]}}
    if speed:
        body["generation_config"] = {"speed": round(speed, 2)}
    res = run("/tts-1.0/generate", body)
    return res["words"][-1]["end"], res

TARGET = 29.5
secs, res = length(open("script.txt").read())
speed = max(0.6, min(1.5, secs / TARGET))
secs, res = length(open("script.txt").read(), speed)
print(speed, secs, res["audio_url"])

Know the edges

  • Speed changes more than length. Past about 1.25 many listeners hear it as rushed, so cut words when the ratio is above that.
  • Speed is not exactly linear. The second take can miss by a few tenths of a second, so check it and nudge once.
  • The end of the last word is not the same as file length. Silence at the tail can add a little, which is why the target leaves half a second.
  • The final loudness is unaffected by speed. Match levels separately with volume.

When the ratio is too big

If the script needs speed above 1.5 or below 0.6, the request is rejected by validation before a job is queued, so you do not pay for it. Cut the copy instead. Keep the one sentence that carries the offer, and the line that the rules require.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume