On-hold phone messages: 12 prompts of 300 characters, cost

Twelve on-hold prompts of 300 characters are 3,600 characters: $0.079 on MAI-Voice-2.1, $0.054 on Flash, and $0.24 on Sume as 12 jobs.

3 min readSume
All posts

Twelve on-hold phone prompts of 300 characters each are 3,600 characters, which cost $0.079 on MAI-Voice-2.1, $0.054 on Flash and $0.24 on Sume TTS as twelve jobs ($0.18 as one).

A phone system usually wants a specific audio format, so the output setting matters as much as the price.

The arithmetic

Rates come from Microsoft's launch page for the two MAI voices and from Sume's public TTS rate. The Sume rate is the list-price arithmetic ($0.0475 per 1,000 characters); the live figure for your workspace is in GET /v1/catalog, and a platform fee can apply on top of list arithmetic on your invoice.

12 prompts of 300 characters: cost by model, rates as of 2026-10-08
OptionRateArithmeticCost
MAI-Voice-2.1$22 per 1M characters3,600 x $22 / 1,000,000$0.079
MAI-Voice-2.1-Flash$15 per 1M characters3,600 x $15 / 1,000,000$0.054
Sume TTS 1.0, 12 separate jobs$0.0475 per 1,000 characters, rounded up to the cent per job12 x ceil(300 x $0.0000475 = $0.014 to the cent) = 12 x $0.02$0.24
Sume TTS 1.0, packed into 1 jobsame rate, at most 20,000 characters per request3,600 characters in 1 request$0.18

Format for a phone line

Sume defaults to mp3 at 44,100 Hz and 128 kbps. If your phone system wants wav, set output_format to wav with pcm_s16le and a sample rate your system accepts; check the sample rates the schema lists before you rely on a specific one.

At 300 characters each, a prompt is 2 cents on Sume (300 x $0.0000475 = $0.01425, rounded up).

  • A speed multiplier below 1 gives a slower, clearer read for hold messages; the range is 0.6 to 1.5.
  • Re-voice a prompt when an opening time changes; that is a 2-cent job.

What Sume does and does not offer here

Sume does not list MAI-Voice-2.1 or MAI-Voice-2.1-Flash, so the first two rows are a price reference, not something you can call through Sume. Sume TTS 1.0 is a managed voice or avatar surface at POST /v1/tts-1.0/generate, and the TTS Router lists Sonic model ids only. A request takes up to 20,000 characters; synthesized audio over 1,200 seconds fails with tts_duration_exceeded and no credit is captured. Microsoft states the Flash model generates up to 45 seconds of audio per call, which matters for long reads.

  • Many short items: one request per item rounds each up to the cent, so packing items into fewer requests (up to 20,000 characters) and slicing by sentence with timestamps.words and segmentation.mode sentence on wav output avoids paying the rounding many times.
  • Output defaults to mp3 at 44,100 Hz and 128 kbps; ask for wav (pcm_s16le, 44,100) when the file feeds a render or a lip-sync step.
  • Speed is a multiplier from 0.6 to 1.5 and volume from 0.5 to 2.0 in generation_config; they do not change the per-character price.

Submitting it on Sume

Paid MCP tools need an idempotency_key, and a cost-only dry run is available before you spend. Use the voice or avatar you already have, set language for any non-English text, and keep one key per item so a retry never bills twice. The tool names are tts_create for synthesis and tts_source_get for the accepted-script manifest.

Treat the table as list arithmetic. Re-check the live rate in the catalog before a large batch, because Microsoft's MAI pricing is a launch announcement dated 2026-10-01 and Sume rates are read from the repository price book as of 2026-10-08.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume