A news app's Listen button: 5,000 backlog articles, voice cost

Voicing a backlog of 5,000 articles at 4,000 characters is 20 million characters: $440.00 on MAI-Voice-2.1, $300.00 on Flash, $950.00 on Sume TTS.

3 min readSume
All posts

Adding a Listen button to a backlog of 5,000 articles of 4,000 characters is 20,000,000 characters, which costs $440.00 on MAI-Voice-2.1, $300.00 on Flash and $950.00 on Sume TTS.

This is a one-time backfill number. New articles after launch are a monthly flow, so the backfill is the part to budget separately.

The arithmetic

Rates come from Microsoft's launch page for the two MAI voices and from Sume's public TTS rate. The Sume rate is the list-price arithmetic ($0.0475 per 1,000 characters); the live figure for your workspace is in GET /v1/catalog, and a platform fee can apply on top of list arithmetic on your invoice.

5,000 articles of 4,000 characters: cost by model, rates as of 2026-10-08
OptionRateArithmeticCost
MAI-Voice-2.1$22 per 1M characters20,000,000 x $22 / 1,000,000$440.00
MAI-Voice-2.1-Flash$15 per 1M characters20,000,000 x $15 / 1,000,000$300.00
Sume TTS 1.0, 5000 separate jobs$0.0475 per 1,000 characters, rounded up to the cent per job5000 x ceil(4000 x $0.0000475 = $0.19 to the cent) = 5000 x $0.19$950.00
Sume TTS 1.0, packed into 1000 jobssame rate, at most 20,000 characters per request20,000,000 characters in 1000 requests$950.00

At this scale the gap is the decision

The Sume figure is 2.16 times the MAI-Voice-2.1 figure ($47.50 divided by $22 per million characters), and 3.17 times the Flash figure. On a 20-million-character backfill that is a gap of hundreds of dollars, not cents.

The gap buys something different: Sume TTS is an asynchronous job surface with idempotent retries and durable hosted files, not a realtime streaming voice. If the Listen button plays audio on tap, you pre-generate and cache the file, so latency does not matter.

  • Pre-generate on publish, store the file, and play the stored copy.
  • At about 15 characters per second (a planning assumption, not a Sume figure), a 4,000-character article is about four to five minutes of speech, well inside the 1,200-second audio cap.

What Sume does and does not offer here

Sume does not list MAI-Voice-2.1 or MAI-Voice-2.1-Flash, so the first two rows are a price reference, not something you can call through Sume. Sume TTS 1.0 is a managed voice or avatar surface at POST /v1/tts-1.0/generate, and the TTS Router lists Sonic model ids only. A request takes up to 20,000 characters; synthesized audio over 1,200 seconds fails with tts_duration_exceeded and no credit is captured. Microsoft states the Flash model generates up to 45 seconds of audio per call, which matters for long reads.

  • Long scripts: split on paragraph boundaries so each request stays at or under 20,000 characters and each audio file under 1,200 seconds.
  • Output defaults to mp3 at 44,100 Hz and 128 kbps; ask for wav (pcm_s16le, 44,100) when the file feeds a render or a lip-sync step.
  • Speed is a multiplier from 0.6 to 1.5 and volume from 0.5 to 2.0 in generation_config; they do not change the per-character price.

Submitting it on Sume

Paid MCP tools need an idempotency_key, and a cost-only dry run is available before you spend. Use the voice or avatar you already have, set language for any non-English text, and keep one key per item so a retry never bills twice. The tool names are tts_create for synthesis and tts_source_get for the accepted-script manifest.

Treat the table as list arithmetic. Re-check the live rate in the catalog before a large batch, because Microsoft's MAI pricing is a launch announcement dated 2026-10-01 and Sume rates are read from the repository price book as of 2026-10-08.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume