Station announcements: 50 messages at 160 characters, Flash vs Sume

Fifty pre-recorded announcements of 160 characters are 8,000 characters: 12 cents on MAI-Voice-2.1-Flash, 18 cents on MAI-Voice-2.1 and 38 cents on Sume.

5 min readSume
All posts

Fifty pre-recorded announcements of 160 characters each are 8,000 characters, costing $0.12 on MAI-Voice-2.1-Flash, $0.18 on MAI-Voice-2.1 and $0.38 on Sume TTS 1.0. At list rates that is 8,000 characters: $0.176 on MAI-Voice-2.1 ($22 per 1M characters), $0.120 on MAI-Voice-2.1-Flash ($15 per 1M), and $0.380 on Sume TTS 1.0 ($47.50 per 1M, or $0.0475 per 1,000 characters).

The inputs are assumptions for this scenario: 50 fixed announcements of 160 characters, generated once and replayed. Microsoft's prices are from its announcement page and Sume's from the API reference and pricing page, all read 2026-10-05. The totals are rate times characters, before any per-job rounding.

What the bill looks like

Fixed messages are generated once, so replays are free of per-character cost on either service.

Cost of 8,000 characters, and of 8,000 characters over one message set, at list rates (read 2026-10-05)
ModelRate per 1M charactersOne runTotal over one message set
MAI-Voice-2.1$22.00$0.176$0.176
MAI-Voice-2.1-Flash$15.00$0.120$0.120
Sume TTS 1.0$47.50$0.380$0.380

Pre-render fixed lines; keep live ones separate

Microsoft's model page lists call-center agents, voice assistants and IVR as the use cases for Flash. A fixed announcement set is different: you know every line in advance, so you can render each once and store the files.

  • Sume jobs are asynchronous with signed webhooks for terminal events, so a batch render needs no open connection.
  • Output WAV at 8,000 or 16,000 Hz where a phone or PA system needs it; Sume's output_format.sample_rate accepts 8000, 16000, 22050, 24000, 44100 and 48000.
  • Use pcm_mulaw or pcm_alaw encoding with a WAV container for telephony.
  • Dynamic lines, such as a platform number read live, are better served by a low-latency model than by a stored file.

Run it on Sume

This renders one fixed announcement as an 8 kHz mu-law WAV.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def speak(text, key, **extra):
    body = {"transcript": text, "language": "en", **extra}
    body.setdefault("avatar_handle", os.environ.get("SUME_AVATAR_HANDLE", ""))
    r = requests.post(f"{API}/v1/tts-1.0/generate", timeout=60,
                      headers={**H, "Idempotency-Key": key}, json=body)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(s["next_poll_after_seconds"] or 2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    res = requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
    return res

res = speak("The next train to the airport leaves from platform four.", "ann-017",
            output_format={"container": "wav", "encoding": "pcm_mulaw", "sample_rate": 8000})
print(res["artifacts"][0]["url"])

When MAI is the better pick

If an announcement must include live data, MAI-Voice-2.1-Flash at about 45 ms (per Microsoft's model page) is the model aimed at that, and Sume TTS is not. For a fixed set rendered once, either works and the price difference is 26 cents.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume