Pre-render a voice line bank for stream alerts instead of live TTS

Alerts, kiosks and game barks repeat a few lines. Render each once with Sume TTS and play the stored file. Thirty 80-character lines cost about 11 cents.

4 min readSume
All posts

If your product speaks a fixed set of lines, such as stream alerts, kiosk prompts or in-game barks, render each line once and play the stored file. A bank costs a few cents, plays with no network wait and sounds identical every time. Sume TTS is billed per character at $0.0475 per 1,000, so 30 lines of 80 characters is 2,400 characters, or $0.114 (API reference).

The trend pulls the other way. Microsoft's MAI-Voice-2.1-Flash is pitched at 150 ms end to end for 45 seconds of audio, at $15 per million characters (Microsoft AI, read 2026-10-04). Fast synthesis is the right tool for text that changes per user. For a fixed line it is spending for speed you do not need.

Where Sume fits and where it does not

A Sume sync call waits at most 30 seconds, and the docs call it a bounded wait, not a streaming channel (jobs docs). It is not built for a voice that answers a viewer in under a second. It is a good fit for a library you build ahead of time.

Build the bank

Give each line a stable name and use that name in the idempotency key. A rerun then returns the same job instead of billing the line again. The script below writes a manifest of name to audio_url.

import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def run(path, body, key):
    r = requests.post(API + path, json=body, timeout=60,
                      headers={**H, "Idempotency-Key": key})
    r.raise_for_status()
    job = r.json()["request_id"]
    while True:
        s = requests.get(f"{API}/v1/jobs/{job}/status", headers=H, timeout=30).json()
        if s.get("terminal"):
            break
        time.sleep(s.get("next_poll_after_seconds") or 3)
    res = requests.get(f"{API}/v1/jobs/{job}/result", headers=H, timeout=30)
    res.raise_for_status()
    return res.json()

LINES = {"sub": "Thank you for the subscription!",
         "raid": "A raid is incoming, say hello!"}
manifest = {}
for name, text in LINES.items():
    res = run("/v1/tts-1.0/generate", {
        "transcript": text,
        "avatar_handle": os.environ["VOICE_HANDLE"],
        "output_format": {"container": "wav", "encoding": "pcm_s16le",
                          "sample_rate": 44100},
    }, f"alert-{name}-v1")
    manifest[name] = res.get("audio_url")
print(manifest)

Keep it maintainable

Version the key suffix (-v1, -v2) when you change a line, so the new text is a new request. Download the files to your own storage if the player lives outside Sume. Check a line with timestamps: {words: true} if you need to cut it to a fixed length. The boundary guide explains the sentence cut.

When to go live instead

If more than a fraction of your lines carry a user name or a number, render live and accept the wait. A mixed design works: pre-render the stable halves, generate the variable part, and join them with timeline audio at $0.01 per join.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume