Pre-render voice agent greetings and fixed replies as TTS files
Live TTS latency is wasted on lines that never change. Render the greeting, hold message and confirmations once as files with Sume TTS and play them instantly.

Lines your voice agent says on every call, such as the greeting, the recording notice, a hold message and a goodbye, do not need a live text-to-speech stage. Render each once as an audio file, store the URL, and play the file when needed. The latency you were trying to shave with a faster synthesiser (Microsoft lists MAI-Voice-2.1-Flash at 150 ms end to end for 45 seconds of audio) drops to the time it takes to start playing a file.
Sume TTS 1.0 is a good fit for the rendering step because it is an async job that returns a hosted file; it is a bad fit for the live reply, because it never streams. This post shows the split.
Which lines are worth pre-rendering?
Pre-rendering pays off when a line is fixed and heard often. The cut line is whether any word changes per call.
- Greeting and recording or consent notice.
- Hold, transfer and timeout messages.
- Confirmation templates with no variables, such as "Your request has been received."
- Menu prompts and error prompts.
- Voicemail drop messages.
| Line type | Render ahead with a Sume job? | Why |
|---|---|---|
| Greeting, hold, goodbye | Yes | Text is fixed; a job result is a durable audio URL |
| Fixed confirmation | Yes | No variables, so one file serves every call |
| Answer that quotes the caller's account | No | Text is unknown until the call; Sume does not stream |
| Number or name spliced into a sentence | Maybe | Render the name or number as its own clip, then join clips with Timeline audio (up to 20 parts) |
How do you render them in one script?
The script submits one TTS job per line with the same avatar voice, waits for each result, and saves the URL in a dictionary you can load from your voice agent at start-up. Output defaults to mp3 at 44.1 kHz and 128 kbps; if your telephony stack wants wav, pass output_format with container wav and encoding pcm_s16le as in the OpenAPI example.
import os, time, requests
BASE, H = "https://api.sume.com", {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = dict(H, **({"Idempotency-Key": key} if key else {}))
d = requests.post(BASE + path, headers=h, json=body, timeout=60)
d.raise_for_status()
d = d.json()["data"]
while not d["terminal"]:
time.sleep(d.get("next_poll_after_seconds") or 2)
d = requests.get(d["status_url"], headers=H, timeout=30).json()["data"]
r = requests.get(d["result_url"], headers=H, timeout=30)
r.raise_for_status() # failed or canceled jobs answer 409 here
return r.json()["data"]["result"]
LINES = {
"greeting": "Thanks for calling. This call may be recorded.",
"hold": "One moment while I look that up.",
"goodbye": "Thanks for calling. Goodbye.",
}
urls = {}
for key, text in LINES.items():
result = run("/v1/tts-1.0/generate", {
"transcript": text,
"avatar_id": os.environ["SUME_AVATAR_ID"],
"mode": "async",
})
urls[key] = result["artifacts"][0]["url"]
print(urls)
How do you keep the files honest?
A synthesised file can skip or change a word, and a fixed line is heard thousands of times, so check each once. Transcribe the rendered file with Sume STT and compare the words with the script; the round-trip check is a short script for it. Version your lines: when the text changes, render again and swap the URL, instead of editing the old file.
If a Sume webhook suits your stack better than polling, mode: webhook with a public HTTPS webhook_url delivers one terminal callback per job; keep status polling as a backup, as the OpenAPI notes.
How big is the saving?
The saving is the difference between synthesising a line live and starting a stored file. Even at a vendor's best claim, such as 150 ms for MAI-Voice-2.1-Flash, a live synthesis stage is network time plus model time on every call, while a file starts as fast as your player can fetch it. On a phone line, where the greeting is the first impression, that is the line to move first.
There is a cost angle too. A fixed 120-character greeting rendered once is one TTS job, about $0.01 at the listed per-character rate rounded up to a cent, instead of one synthesis per call. At a thousand calls a day the live version repeats that thousand times for the same audio.
Keep the lines in version control next to the script that renders them, with a hash of each text, so a change to the words triggers a re-render and a stale file never plays.
Where does the live path come from?
The remainder of the call needs a streaming stack. Microsoft's MAI-Transcribe-2-Streaming reports first partials in just over 100 ms, and Inception's Mercury Voice lists a median under 320 ms to the first answer token; both are vendor claims from their own pages, read 2026-10-03. Your speech and language stages handle the unpredictable replies, and the files handle everything you can write down in advance.
Sources
Related posts
More in Use cases
- Recast one video ad with several faces: consent checklist, cost
Reuse one 15 second ad with different people via H3 Max Recast on Sume: cost per variant, the release to get from each person, and how to label them.
- Recast a video with a shot over 15 seconds: split it first on Sume
H3 Max Recast takes 5-30 s with no shot over 15 s. How to trim a long take, keep the same people on each piece, and join the results in a Timeline.
- Recover YouTube Shorts reach after reposting other creators' clips
YouTube says a channel that shifts from re-uploading to original content gets its reach re-evaluated. What is known, what isn't, and a plan using Sume.
- Replace someone with yourself in a video with AI: Recast on Sume
Put yourself into an existing clip: send a public source video and one photo of you to H3 Max Recast on Sume. Limits, cost per clip and consent rules.
Written by Sume