Is MAI-Voice-2.1 on Sume? What Sume's audio docs list instead

Sume does not list MAI-Voice-2.1 or MAI-Transcribe-2. Here is what its docs do list for speech, transcripts, music and captions, with prices.

5 min readSume
All posts

No. Sume's model docs do not list MAI-Voice-2.1, MAI-Voice-2.1-Flash or MAI-Transcribe-2-Streaming, so you cannot call them through Sume. What Sume lists for audio work is a set of its own surfaces: text-to-speech and speech-to-text tools, Music through the Music Router, burned-in captions, audio detach and timeline audio. This post sets the Microsoft announcement next to those surfaces so you can see what overlaps and what does not.

What Microsoft announced

The facts below come from Microsoft AI's own announcement page, read 2026-10-08. They describe Microsoft's products, not anything Sume ships.

Microsoft's published figures, read 2026-10-08
ModelPublished figure
MAI-Voice-2.1$22 per 1M characters; 23 languages and 26 locales
MAI-Voice-2.1-Flash$15 per 1M characters; end-to-end latency 150 ms
MAI-Transcribe-2-Streaming$0.54 per hour of audio through the end of the year; 60 languages with automatic detection
Voice cloning (Microsoft)From a few seconds of reference audio

What Sume lists in the same area

The Sume figures below come from the Sume docs pages named in the sources. None of them is a Microsoft model.

Sume audio surfaces, as of 2026-10-08
NeedSume surfaceDocumented price
Speech to text on a clipVideo inspect with transcribe: true$0.01 per audio minute (plus inspect compute)
MusicMusic Router, sume/music-auto$0.125 per accepted generation
Burned-in captionsVideo captions$0.20 per job, clips up to 60 s
Audio track as wav or mp3Audio detach$0.01 per job
Join or split audioTimeline audio$0.01 flat per job

What the docs do not say

The Sume docs pages I read do not describe a voice-cloning route, so this post makes no cloning comparison and does not claim that Sume matches Microsoft's reference-audio feature. The MCP page lists tts_create and stt_create as tools, but I did not use it to quote a TTS price here; check GET /v1/catalog for the live TTS rate before you plan a budget.

Streaming is also not a Sume feature in the docs read. Sume's transcript path takes a stored clip, runs a job, and returns the transcript with optional sentence segments. Microsoft's streaming model targets live conversation, which is a different workload.

How to decide

If you need live voice-agent latency, the Microsoft figures above are the relevant ones and you would use Microsoft's own endpoints. If you need to turn finished clips into transcripts, captions, scored video and joined narration inside one workflow, the Sume surfaces cover it. Many teams need both, and they do not conflict.

  • Check the vendor page for current pricing; the $0.54 per hour rate is stated as introductory through the end of the year.
  • Check GET /v1/catalog for Sume's live rates.
  • Sume does not list MAI models; do not assume a model id exists.

Sources

Related posts

More in Models

All Models posts

Written by Sume