Pre-render voice agent greetings and fixed replies as TTS files

Live TTS latency is wasted on lines that never change. Render the greeting, hold message and confirmations once as files with Sume TTS and play them instantly.

5 min readSume
All posts

Lines your voice agent says on every call, such as the greeting, the recording notice, a hold message and a goodbye, do not need a live text-to-speech stage. Render each once as an audio file, store the URL, and play the file when needed. The latency you were trying to shave with a faster synthesiser (Microsoft lists MAI-Voice-2.1-Flash at 150 ms end to end for 45 seconds of audio) drops to the time it takes to start playing a file.

Sume TTS 1.0 is a good fit for the rendering step because it is an async job that returns a hosted file; it is a bad fit for the live reply, because it never streams. This post shows the split.

Which lines are worth pre-rendering?

Pre-rendering pays off when a line is fixed and heard often. The cut line is whether any word changes per call.

  • Greeting and recording or consent notice.
  • Hold, transfer and timeout messages.
  • Confirmation templates with no variables, such as "Your request has been received."
  • Menu prompts and error prompts.
  • Voicemail drop messages.
What to render ahead and what to keep live, from the Sume TTS request schema, read 2026-10-03.
Line typeRender ahead with a Sume job?Why
Greeting, hold, goodbyeYesText is fixed; a job result is a durable audio URL
Fixed confirmationYesNo variables, so one file serves every call
Answer that quotes the caller's accountNoText is unknown until the call; Sume does not stream
Number or name spliced into a sentenceMaybeRender the name or number as its own clip, then join clips with Timeline audio (up to 20 parts)

How do you render them in one script?

The script submits one TTS job per line with the same avatar voice, waits for each result, and saves the URL in a dictionary you can load from your voice agent at start-up. Output defaults to mp3 at 44.1 kHz and 128 kbps; if your telephony stack wants wav, pass output_format with container wav and encoding pcm_s16le as in the OpenAPI example.

import os, time, requests
BASE, H = "https://api.sume.com", {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
    h = dict(H, **({"Idempotency-Key": key} if key else {}))
    d = requests.post(BASE + path, headers=h, json=body, timeout=60)
    d.raise_for_status()
    d = d.json()["data"]
    while not d["terminal"]:
        time.sleep(d.get("next_poll_after_seconds") or 2)
        d = requests.get(d["status_url"], headers=H, timeout=30).json()["data"]
    r = requests.get(d["result_url"], headers=H, timeout=30)
    r.raise_for_status()  # failed or canceled jobs answer 409 here
    return r.json()["data"]["result"]

LINES = {
    "greeting": "Thanks for calling. This call may be recorded.",
    "hold": "One moment while I look that up.",
    "goodbye": "Thanks for calling. Goodbye.",
}
urls = {}
for key, text in LINES.items():
    result = run("/v1/tts-1.0/generate", {
        "transcript": text,
        "avatar_id": os.environ["SUME_AVATAR_ID"],
        "mode": "async",
    })
    urls[key] = result["artifacts"][0]["url"]
print(urls)

How do you keep the files honest?

A synthesised file can skip or change a word, and a fixed line is heard thousands of times, so check each once. Transcribe the rendered file with Sume STT and compare the words with the script; the round-trip check is a short script for it. Version your lines: when the text changes, render again and swap the URL, instead of editing the old file.

If a Sume webhook suits your stack better than polling, mode: webhook with a public HTTPS webhook_url delivers one terminal callback per job; keep status polling as a backup, as the OpenAPI notes.

How big is the saving?

The saving is the difference between synthesising a line live and starting a stored file. Even at a vendor's best claim, such as 150 ms for MAI-Voice-2.1-Flash, a live synthesis stage is network time plus model time on every call, while a file starts as fast as your player can fetch it. On a phone line, where the greeting is the first impression, that is the line to move first.

There is a cost angle too. A fixed 120-character greeting rendered once is one TTS job, about $0.01 at the listed per-character rate rounded up to a cent, instead of one synthesis per call. At a thousand calls a day the live version repeats that thousand times for the same audio.

Keep the lines in version control next to the script that renders them, with a hash of each text, so a change to the words triggers a re-render and a stale file never plays.

Where does the live path come from?

The remainder of the call needs a streaming stack. Microsoft's MAI-Transcribe-2-Streaming reports first partials in just over 100 ms, and Inception's Mercury Voice lists a median under 320 ms to the first answer token; both are vendor claims from their own pages, read 2026-10-03. Your speech and language stages handle the unpredictable replies, and the files handle everything you can write down in advance.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume