Onboarding voice tips: 25 screens, 3 narrators, MAI vs Sume cost

Testing three narrators on 25 onboarding screens of 120 characters is 9,000 characters: 14 cents on MAI Flash, 20 cents on MAI-Voice-2.1, 43 cents on Sume.

5 min readSume
All posts

Trying three narrators on 25 onboarding screens of 120 characters each is 9,000 characters, which costs $0.14 on MAI-Voice-2.1-Flash, $0.20 on MAI-Voice-2.1 and $0.43 on Sume TTS 1.0. At list rates that is 9,000 characters: $0.198 on MAI-Voice-2.1 ($22 per 1M characters), $0.135 on MAI-Voice-2.1-Flash ($15 per 1M), and $0.427 on Sume TTS 1.0 ($47.50 per 1M, or $0.0475 per 1,000 characters).

The inputs are assumptions for this scenario: 25 screens, 120 characters each, three voices, one pass each. Microsoft's prices are from its announcement page and Sume's from the API reference and pricing page, all read 2026-10-05. The totals are rate times characters, before any per-job rounding.

What the bill looks like

The cost of testing voices is tiny; the cost that matters is the time to compare them, so generate all three sets in one go.

Cost of 9,000 characters, and of 9,000 characters over one three-voice test, at list rates (read 2026-10-05)
ModelRate per 1M charactersOne runTotal over one three-voice test
MAI-Voice-2.1$22.00$0.198$0.198
MAI-Voice-2.1-Flash$15.00$0.135$0.135
Sume TTS 1.0$47.50$0.427$0.427

Hold everything but the voice constant

An A/B test of narrators is only fair if the text, speed, volume and language are identical. Keep one config and change only the voice selector.

  • On Sume the voice is chosen with avatar_handle, avatar_id or voice.id; the rest of the request can stay byte-identical.
  • Use one idempotency key per voice and screen, for example onb-voiceB-screen07, so a rerun does not double-bill.
  • Microsoft lists instant voice matching and granular emotion control for MAI-Voice-2.1, so a matched voice can be one of your candidates there.
  • Language mismatches are caught before a job is created: a voice whose language differs from language returns 409 tts_voice_language_mismatch with no charge.

Run it on Sume

This renders the same 25 lines with three voice selectors and prints a URL per voice per screen. Set the three handles in the environment.

import os, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def speak(text, key, **extra):
    body = {"transcript": text, "language": "en", **extra}
    body.setdefault("avatar_handle", os.environ.get("SUME_AVATAR_HANDLE", ""))
    r = requests.post(f"{API}/v1/tts-1.0/generate", timeout=60,
                      headers={**H, "Idempotency-Key": key}, json=body)
    r.raise_for_status()
    job = r.json()["data"]
    while True:
        s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
        if s["terminal"]:
            break
        time.sleep(s["next_poll_after_seconds"] or 2)
    if s["sume_status"] != "completed":
        raise RuntimeError(s["sume_status"])
    res = requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
    return res

handles = os.environ["SUME_AVATAR_HANDLES"].split(",")   # three handles
screens = ["Tap the plus to add your first item.", "Swipe left to archive."]
for h, handle in enumerate(handles):
    for i, line in enumerate(screens, 1):
        res = speak(line, f"onb-v{h}-s{i:02d}", avatar_handle=handle)
        print(h, i, res["artifacts"][0]["url"])

When MAI is the better pick

If the final product will speak in a live session, a low-latency model is the right test and Microsoft's Flash page is aimed at it. For pre-rendered tips, price differences of cents do not decide this; pick the voice that tests best.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume