MAI-Voice-2.1 pricing vs Sume TTS: what Sume lists, what it does not

MAI-Voice-2.1 costs $22 per 1M characters per Microsoft AI. Sume does not list MAI; its TTS runs Sonic voices at $47.50 per 1M characters.

5 min readSume
All posts

Sume does not list Microsoft's MAI-Voice-2.1 or MAI-Transcribe-2-Streaming; its text-to-speech route runs Sonic voices and its speech-to-text route is a separate fixed-flag job. Microsoft AI prices MAI-Voice-2.1 at $22 per 1M characters, against Sume TTS at $47.50 per 1M characters ($0.0475 per 1,000), so if you want MAI today you buy it where Microsoft sells it.

Microsoft facts are from its page Our first streaming transcription model debuts at no. 1 on Artificial Analysis (read 2026-10-10). Sume facts are from the Sume OpenAPI contract and the TTS router catalog in the repository.

What Microsoft says it launched

The page names MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. It says MAI-Voice-2.1 supports 23 languages and 26 locales, and that Flash generates 45 seconds of audio with 150 ms end-to-end latency. It lists availability through Microsoft Foundry, OpenRouter, Vercel and the MAI Playground, with LiveKit coming soon. The transcription model is priced at $0.54 per hour as an introductory rate through year-end.

Prices as published (Microsoft AI read 2026-10-10; Sume per golden pricing tests)
ItemPriceUnit
MAI-Voice-2.1$22per 1M characters
MAI-Voice-2.1-Flash$15per 1M characters
MAI-Transcribe-2-Streaming$0.54 (introductory)per hour
Sume TTS 1.0$47.50 ($0.0475 per 1,000)per 1M characters
Sume STT 1.0$0.60 ($0.01 per minute)per hour

What Sume lists instead

The TTS router accepts sonic-3.6, sonic-3.5, sonic-3, latest (an alias for 3.6) and preview, which is a beta that returns voice_model_mismatch for professional clones. Sume STT takes a public audio URL of up to 600 seconds, returns word timings, and is billed at $0.01 per audio minute. It does not do streaming; a job is a request and a result.

That last point is the real difference with MAI-Transcribe-2-Streaming. A streaming model returns partial text while the speaker is still talking. Sume's job envelope has a sync wait of up to 30 seconds, but that is waiting for one finished result, not a live stream.

How to decide

The price gap is real, but price per character is only part of a voice project. Compare the languages you need, whether your voice has to be a specific person, and whether you need live partials.

  • Live captions on a call: a streaming model is the right tool; Sume STT is not.
  • Voice-over for finished videos: Sume's TTS plus timeline audio keeps it in one project, and it is $0.0475 per 1,000 characters.
  • A specific cloned voice: cloning is app-only on Sume; the API takes a voice id you already have.
  • Lowest cost per character: on published prices, the MAI figures are lower.

A caution on the Microsoft numbers

The $0.54 transcription price is described as introductory through year-end, so it may change in January. The latency and accuracy claims are Microsoft's own statements on its own page; I did not measure them and the comparison above does not use them.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume