Is MAI-Voice-2.1 on Sume? What Sume's audio docs list instead
Sume does not list MAI-Voice-2.1 or MAI-Transcribe-2. Here is what its docs do list for speech, transcripts, music and captions, with prices.

No. Sume's model docs do not list MAI-Voice-2.1, MAI-Voice-2.1-Flash or MAI-Transcribe-2-Streaming, so you cannot call them through Sume. What Sume lists for audio work is a set of its own surfaces: text-to-speech and speech-to-text tools, Music through the Music Router, burned-in captions, audio detach and timeline audio. This post sets the Microsoft announcement next to those surfaces so you can see what overlaps and what does not.
What Microsoft announced
The facts below come from Microsoft AI's own announcement page, read 2026-10-08. They describe Microsoft's products, not anything Sume ships.
| Model | Published figure |
|---|---|
| MAI-Voice-2.1 | $22 per 1M characters; 23 languages and 26 locales |
| MAI-Voice-2.1-Flash | $15 per 1M characters; end-to-end latency 150 ms |
| MAI-Transcribe-2-Streaming | $0.54 per hour of audio through the end of the year; 60 languages with automatic detection |
| Voice cloning (Microsoft) | From a few seconds of reference audio |
What Sume lists in the same area
The Sume figures below come from the Sume docs pages named in the sources. None of them is a Microsoft model.
| Need | Sume surface | Documented price |
|---|---|---|
| Speech to text on a clip | Video inspect with transcribe: true | $0.01 per audio minute (plus inspect compute) |
| Music | Music Router, sume/music-auto | $0.125 per accepted generation |
| Burned-in captions | Video captions | $0.20 per job, clips up to 60 s |
| Audio track as wav or mp3 | Audio detach | $0.01 per job |
| Join or split audio | Timeline audio | $0.01 flat per job |
What the docs do not say
The Sume docs pages I read do not describe a voice-cloning route, so this post makes no cloning comparison and does not claim that Sume matches Microsoft's reference-audio feature. The MCP page lists tts_create and stt_create as tools, but I did not use it to quote a TTS price here; check GET /v1/catalog for the live TTS rate before you plan a budget.
Streaming is also not a Sume feature in the docs read. Sume's transcript path takes a stored clip, runs a job, and returns the transcript with optional sentence segments. Microsoft's streaming model targets live conversation, which is a different workload.
How to decide
If you need live voice-agent latency, the Microsoft figures above are the relevant ones and you would use Microsoft's own endpoints. If you need to turn finished clips into transcripts, captions, scored video and joined narration inside one workflow, the Sume surfaces cover it. Many teams need both, and they do not conflict.
- Check the vendor page for current pricing; the $0.54 per hour rate is stated as introductory through the end of the year.
- Check
GET /v1/catalogfor Sume's live rates. - Sume does not list MAI models; do not assume a model id exists.
Sources
Related posts
More in Models
- Kling 3 audio on vs off for 100 fifteen-second SKU ads: $105 gap
Kling 3 Pro on Sume is $0.14 per second with audio off and $0.21 on. At 15 s across 100 SKUs that is $210 vs $315, a $105 difference.
- Kling 3 stops at 15 seconds: three ways to a 30-second clip
Kling 3 on Sume takes 4 to 15 seconds. For 30 seconds, use Wan 3.0 at $3.75 (720p) or Seedance 2.5, or join two takes with Timeline. Prices side by side.
- Kling 4.0 API: what to generate with on Sume while you wait
Kling 4.0 was announced in September 2026. Sume's catalog lists kling-3 today: 4 to 15 seconds, $0.14 a second silent, $0.21 with sound. Request included.
- Kling.ai headlines Kling 4.0: what an API buyer can call today
kling.ai now shows 'All-New Kling 4.0', but fal still says full launch is October. Sume lists kling-3 (4-15 s, 720p/1080p) and no 4.0 id. Prices inside.
Written by Sume