Compliance training narration: 30 modules at 8,000 characters

30 training modules of 8,000 characters are 240,000 characters: $5.28 on MAI-Voice-2.1, $3.60 on Flash, $11.40 on Sume TTS as 30 jobs.

3 min readSume
All posts

Narrating 30 compliance modules of 8,000 characters each is 240,000 characters, which costs $5.28 on MAI-Voice-2.1, $3.60 on Flash and $11.40 on Sume TTS as 30 jobs.

Compliance content is revised every year, so the real figure is the cost of a full re-voice, not only the first pass.

The arithmetic

Rates come from Microsoft's launch page for the two MAI voices and from Sume's public TTS rate. The Sume rate is the list-price arithmetic ($0.0475 per 1,000 characters); the live figure for your workspace is in GET /v1/catalog, and a platform fee can apply on top of list arithmetic on your invoice.

30 modules of 8,000 characters: cost by model, rates as of 2026-10-08
OptionRateArithmeticCost
MAI-Voice-2.1$22 per 1M characters240,000 x $22 / 1,000,000$5.28
MAI-Voice-2.1-Flash$15 per 1M characters240,000 x $15 / 1,000,000$3.60
Sume TTS 1.0, 30 separate jobs$0.0475 per 1,000 characters, rounded up to the cent per job30 x ceil(8000 x $0.0000475 = $0.38 to the cent) = 30 x $0.38$11.40
Sume TTS 1.0, packed into 12 jobssame rate, at most 20,000 characters per request240,000 characters in 12 requests$11.40

Budget the annual refresh

Each 8,000-character module is 38 cents on Sume (8,000 x $0.0000475 = $0.38). The full set is $11.40 whether you send 30 jobs or 12 packed requests, because both fill the rounding evenly.

If only 6 modules change in a year, the refresh is $2.28 on Sume. Use the module id plus a revision number as the idempotency key so a retry is safe and a re-voice is a new job.

  • At about 15 characters per second (a planning assumption, not a Sume figure), 8,000 characters is about nine minutes of speech, under the 1,200-second audio cap.
  • Add the cost of captions and the render separately; they are not part of the voice figure.

What Sume does and does not offer here

Sume does not list MAI-Voice-2.1 or MAI-Voice-2.1-Flash, so the first two rows are a price reference, not something you can call through Sume. Sume TTS 1.0 is a managed voice or avatar surface at POST /v1/tts-1.0/generate, and the TTS Router lists Sonic model ids only. A request takes up to 20,000 characters; synthesized audio over 1,200 seconds fails with tts_duration_exceeded and no credit is captured. Microsoft states the Flash model generates up to 45 seconds of audio per call, which matters for long reads.

  • Long scripts: split on paragraph boundaries so each request stays at or under 20,000 characters and each audio file under 1,200 seconds.
  • Output defaults to mp3 at 44,100 Hz and 128 kbps; ask for wav (pcm_s16le, 44,100) when the file feeds a render or a lip-sync step.
  • Speed is a multiplier from 0.6 to 1.5 and volume from 0.5 to 2.0 in generation_config; they do not change the per-character price.

Submitting it on Sume

Paid MCP tools need an idempotency_key, and a cost-only dry run is available before you spend. Use the voice or avatar you already have, set language for any non-English text, and keep one key per item so a retry never bills twice. The tool names are tts_create for synthesis and tts_source_get for the accepted-script manifest.

Treat the table as list arithmetic. Re-check the live rate in the catalog before a large batch, because Microsoft's MAI pricing is a launch announcement dated 2026-10-01 and Sume rates are read from the repository price book as of 2026-10-08.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume