Pre-render a voice line bank for stream alerts instead of live TTS
Alerts, kiosks and game barks repeat a few lines. Render each once with Sume TTS and play the stored file. Thirty 80-character lines cost about 11 cents.

If your product speaks a fixed set of lines, such as stream alerts, kiosk prompts or in-game barks, render each line once and play the stored file. A bank costs a few cents, plays with no network wait and sounds identical every time. Sume TTS is billed per character at $0.0475 per 1,000, so 30 lines of 80 characters is 2,400 characters, or $0.114 (API reference).
The trend pulls the other way. Microsoft's MAI-Voice-2.1-Flash is pitched at 150 ms end to end for 45 seconds of audio, at $15 per million characters (Microsoft AI, read 2026-10-04). Fast synthesis is the right tool for text that changes per user. For a fixed line it is spending for speed you do not need.
Where Sume fits and where it does not
A Sume sync call waits at most 30 seconds, and the docs call it a bounded wait, not a streaming channel (jobs docs). It is not built for a voice that answers a viewer in under a second. It is a good fit for a library you build ahead of time.
Build the bank
Give each line a stable name and use that name in the idempotency key. A rerun then returns the same job instead of billing the line again. The script below writes a manifest of name to audio_url.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def run(path, body, key):
r = requests.post(API + path, json=body, timeout=60,
headers={**H, "Idempotency-Key": key})
r.raise_for_status()
job = r.json()["request_id"]
while True:
s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
if s.get("terminal"):
break
time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
res.raise_for_status()
return res.json()
LINES = {"sub": "Thank you for the subscription!",
"raid": "A raid is incoming, say hello!"}
manifest = {}
for name, text in LINES.items():
res = run("/v1/tts-1.0/generate", {
"transcript": text,
"avatar_handle": os.environ["VOICE_HANDLE"],
"output_format": {"container": "wav", "encoding": "pcm_s16le",
"sample_rate": 44100},
}, f"alert-{name}-v1")
manifest[name] = res.get("audio_url")
print(manifest)Keep it maintainable
Version the key suffix (-v1, -v2) when you change a line, so the new text is a new request. Download the files to your own storage if the player lives outside Sume. Check a line with timestamps: {words: true} if you need to cut it to a fixed length. The boundary guide explains the sentence cut.
When to go live instead
If more than a fraction of your lines carry a user name or a number, render live and accept the wait. A mixed design works: pre-render the stable halves, generate the variable part, and join them with timeline audio at $0.01 per join.
Sources
Related posts
More in Use cases
- Price increase announcement video: a 45-second avatar, four scenes
Announce a price change with a 45-second Sume avatar video in four scenes: hook, reason, what changes, next step. Cost by tier, a working video_inputs body.
- Add print bleed to an AI image: mirror padding in Pillow
An AI image has no bleed, so a trimmed print can show a white edge. Extend the edges by mirroring with NumPy and Pillow, to the bleed your printer specifies.
- Privacy policy update explainer: a 40-second avatar video
Explain a privacy policy change in 40 seconds with a Sume avatar: three scenes, captions on, legal review before render. Cost by tier and a script checklist.
- Pumpkin patch last-weekend promo clip with a carving-night hook
Use the NRF carving stat to make a 10-second last-weekend promo for a pumpkin patch: two Wan 3.0 shots and a Timeline join, about $0.73 at 480p on Sume.
Written by Sume