Museum audio guide: 40 stops at 600 characters, voice cost

A 40-stop audio guide at 600 characters a stop is 24,000 characters: $0.528 on MAI-Voice-2.1, $0.36 on Flash, $1.20 on Sume as 40 jobs.

3 min readSume
All posts

A 40-stop audio guide with 600 characters per stop is 24,000 characters, which costs $0.528 on MAI-Voice-2.1, $0.36 on MAI-Voice-2.1-Flash and $1.20 on Sume TTS as 40 separate jobs.

Guides are re-recorded whenever a label or an exhibit changes, so the cost that matters is the cost of one stop, not just the first pass.

The arithmetic

Rates come from Microsoft's launch page for the two MAI voices and from Sume's public TTS rate. The Sume rate is the list-price arithmetic ($0.0475 per 1,000 characters); the live figure for your workspace is in GET /v1/catalog, and a platform fee can apply on top of list arithmetic on your invoice.

40 stops of 600 characters: cost by model, rates as of 2026-10-08
OptionRateArithmeticCost
MAI-Voice-2.1$22 per 1M characters24,000 x $22 / 1,000,000$0.528
MAI-Voice-2.1-Flash$15 per 1M characters24,000 x $15 / 1,000,000$0.36
Sume TTS 1.0, 40 separate jobs$0.0475 per 1,000 characters, rounded up to the cent per job40 x ceil(600 x $0.0000475 = $0.029 to the cent) = 40 x $0.03$1.20
Sume TTS 1.0, packed into 2 jobssame rate, at most 20,000 characters per request24,000 characters in 2 requests$1.14

Updating one stop

On Sume, one 600-character stop is 3 cents (600 x $0.0000475 = $0.0285, rounded up), so a corrected script costs 3 cents to re-voice. The total for all 40 is $1.20 as separate jobs.

All 40 stops together are 24,000 characters, which does not fit one 20,000-character request. Two requests would cover them, but you would lose per-stop re-runs.

  • Keep the same voice id for every stop so the guide sounds like one narrator.
  • Set language on every non-English stop; a mismatch between voice language and text language returns a 409 warning that you confirm before the job runs.

What Sume does and does not offer here

Sume does not list MAI-Voice-2.1 or MAI-Voice-2.1-Flash, so the first two rows are a price reference, not something you can call through Sume. Sume TTS 1.0 is a managed voice or avatar surface at POST /v1/tts-1.0/generate, and the TTS Router lists Sonic model ids only. A request takes up to 20,000 characters; synthesized audio over 1,200 seconds fails with tts_duration_exceeded and no credit is captured. Microsoft states the Flash model generates up to 45 seconds of audio per call, which matters for long reads.

  • Many short items: one request per item rounds each up to the cent, so packing items into fewer requests (up to 20,000 characters) and slicing by sentence with timestamps.words and segmentation.mode sentence on wav output avoids paying the rounding many times.
  • Output defaults to mp3 at 44,100 Hz and 128 kbps; ask for wav (pcm_s16le, 44,100) when the file feeds a render or a lip-sync step.
  • Speed is a multiplier from 0.6 to 1.5 and volume from 0.5 to 2.0 in generation_config; they do not change the per-character price.

Submitting it on Sume

Paid MCP tools need an idempotency_key, and a cost-only dry run is available before you spend. Use the voice or avatar you already have, set language for any non-English text, and keep one key per item so a retry never bills twice. The tool names are tts_create for synthesis and tts_source_get for the accepted-script manifest.

Treat the table as list arithmetic. Re-check the live rate in the catalog before a large batch, because Microsoft's MAI pricing is a launch announcement dated 2026-10-01 and Sume rates are read from the repository price book as of 2026-10-08.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume