Weekly 6,000-character guided meditation: MAI-Voice-2.1 vs Sume cost
A 6,000-character meditation script a week for a year: $6.86 on MAI-Voice-2.1, $4.68 on Flash, $14.82 on Sume TTS. Plus the 45-second and 1,200-second caps.

A guided meditation of 6,000 characters released every week for a year costs about $6.86 on MAI-Voice-2.1, $4.68 on MAI-Voice-2.1-Flash and $14.82 on Sume TTS 1.0. At list rates that is 312,000 characters: $6.86 on MAI-Voice-2.1 ($22 per 1M characters), $4.68 on MAI-Voice-2.1-Flash ($15 per 1M), and $14.82 on Sume TTS 1.0 ($47.50 per 1M, or $0.0475 per 1,000 characters).
The inputs are assumptions for this scenario: one 6,000-character script per week, 52 weeks, one voice. Meditation pacing varies, so check how long your own audio runs. Microsoft's prices are from its announcement page and Sume's from the API reference and pricing page, all read 2026-10-05. The totals are rate times characters, before any per-job rounding.
What the bill looks like
Long, slow scripts are the best case for per-character pricing: you pay for the text, not for the silence the narrator leaves between lines.
| Model | Rate per 1M characters | One run | Total over 52 weekly episodes |
|---|---|---|---|
| MAI-Voice-2.1 | $22.00 | $0.132 | $6.86 |
| MAI-Voice-2.1-Flash | $15.00 | $0.090 | $4.68 |
| Sume TTS 1.0 | $47.50 | $0.285 | $14.82 |
Length caps matter more than price here
A meditation is one long take, so each service's ceiling decides how you build it. Microsoft's announcement lists a 45-second audio generation capability for MAI-Voice-2.1-Flash, so a long session would need to be cut into pieces on that model. Microsoft's model page lists MAI-Voice-2.1 as the one to use when fidelity matters more than speed, and describes it as holding up across long generations.
- Sume TTS 1.0 takes up to 20,000 characters per request, so a 6,000-character script is a single job.
- Synthesized audio longer than 1,200 seconds fails with
tts_duration_exceededand no credit is captured, so a very slow reading is split by section instead. generation_config.speedaccepts 0.6 to 1.5, which is the control for a slower, calmer pace.- The default output is MP3 at 44.1 kHz and 128 kbps; ask for WAV if you will mix a music bed under it.
Run it on Sume
Send the script as one job at a slower speed and read the audio URL from the result. The job is asynchronous, so poll the status URL rather than holding the request open; if your client times out, retry with the same Idempotency-Key and Sume returns the original job.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def speak(text, key, **extra):
body = {"transcript": text, "language": "en", **extra}
body.setdefault("avatar_handle", os.environ.get("SUME_AVATAR_HANDLE", ""))
r = requests.post(f"{API}/v1/tts-1.0/generate", timeout=60,
headers={**H, "Idempotency-Key": key}, json=body)
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=H, timeout=30).json()["data"]
if s["terminal"]:
break
time.sleep(s["next_poll_after_seconds"] or 2)
if s["sume_status"] != "completed":
raise RuntimeError(s["sume_status"])
res = requests.get(job["result_url"], headers=H, timeout=30).json()["data"]["result"]
return res
script = open("meditation.txt", encoding="utf-8").read()
res = speak(script, "meditation-week-41", generation_config={"speed": 0.85})
print(res["artifacts"][0]["url"])
print(f"estimated {len(script) * 47.5 / 1e6:.3f} USD at list rate")
When MAI is the better pick
Pick MAI-Voice-2.1 when the cheaper per-character rate matters at volume and a Foundry-based pipeline is already in place: at 52 episodes the gap is $7.96 a year on this script. Pick Sume when the voiceover is one step in a chain that also trims, mixes and renders on one API key, and you want a job URL instead of a model deployment.
Sources
Related posts
More in Pricing
- Weekly newsletter audio edition: a year of TTS at 6,000 characters
Reading a 6,000-character newsletter aloud every week is 312,000 characters a year: $4.68 on MAI Flash, $6.86 on MAI-Voice-2.1, $14.82 on Sume TTS.
- What $100 buys in transcription hours: MAI Streaming vs Sume STT
$100 buys about 185 audio hours on MAI-Transcribe-2-Streaming at its $0.54 intro rate, and 166 hours on Sume STT at $0.01 a minute.
- What $50 buys in 5-second AI video clips on each Sume model
$50 buys 131 five-second clips on H3 at 768p or 17 on Seedance 2.5 at 720p on Sume. Ten rows, clip counts rounded down.
- What $50 buys on Sume: speech, transcripts, images, music, video
$50 buys 1.05M characters of speech, 83 hours of transcription, 400 music tracks or 400 seconds of Wan 720p, each at its published rate.
Written by Sume