MAI-Transcribe-2 at $0.54 vs Sume STT at $0.60 an hour: the 6-cent gap
MAI-Transcribe-2-Streaming is $0.54 an audio hour through 2026; Sume STT is $0.60 an hour. At 1,000 hours the gap is $60. They suit different jobs.

MAI-Transcribe-2-Streaming costs $0.54 per hour of audio through the end of 2026 (Microsoft AI, read 2026-10-09). Sume STT 1.0 costs $0.01 per audio minute, which is $0.60 per hour (Sume API catalog, read 2026-10-09). The gap is 6 cents an hour: $6 per 100 hours, $60 per 1,000 hours. Microsoft calls its price introductory, and the page I read states no price for 2027, so only 2026 volume can be compared today.
The numbers
Both rates are linear in audio length, so the comparison is a multiplication.
| Audio per month | MAI-Transcribe-2-Streaming at $0.54/h | Sume STT at $0.60/h | Difference |
|---|---|---|---|
| 100 hours | $54.00 | $60.00 | $6.00 |
| 500 hours | $270.00 | $300.00 | $30.00 |
| 1,000 hours | $540.00 | $600.00 | $60.00 |
| 2,500 hours | $1,350.00 | $1,500.00 | $150.00 |
What each service is built for
Microsoft describes MAI-Transcribe-2-Streaming as a real-time model: it reports first transcripts in just over 100 ms, partial hypotheses while a person is still speaking, 60 languages and continuous language detection. It is listed as available in Microsoft Foundry, the MAI Playground, Vercel and Azure Voice Live, with LiveKit coming soon. That fits live captions and voice agents.
Sume STT 1.0 is a job: you send a public HTTPS audio_url, optionally a language_code hint and a duration_seconds between 1 and 600, then poll the job or take a webhook. The result carries the text and word timings, and a segmentation option adds sentence segments. It is built for finished recordings that feed captions, edits or search, so it is not a competitor for live use.
When the 6 cents does not decide it
If you need words on screen while someone talks, Sume's batch job cannot do it and the price is moot. If you transcribe finished videos, the 6 cents an hour is small next to the steps around it: cutting audio, captioning, rendering. A one-hour recording on Sume costs $0.60 to transcribe, and the same job on MAI would be $0.54; neither changes the cost of the edit around it much.
Consider the shape of the job too. Sume's reserve works from the duration you supply, so give it a measured duration_seconds. A ten-minute ceiling per request means a long recording is cut into pieces first, which the audio detach docs and Timeline audio docs cover. MAI's page lists no such per-request ceiling in the text I read.
A quick way to decide
Write down three facts: whether a person is waiting for the words as they speak, how many hours you transcribe, and whether the recording already sits in an editing workflow. If the first answer is yes, use a streaming service. If it is no, the second fact only changes the bill by $0.06 an hour, and the third decides the rest: a transcript that is already a job with word timings can feed captions without a second upload. Re-check Microsoft's price page before you commit, since the $0.54 rate is stated for the rest of 2026 only.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 lists 23 languages; how Sume TTS sets its language
Microsoft lists 23 languages and 26 locales for MAI-Voice-2.1. Sume TTS takes one language field, defaults to English and asks you to confirm a voice mismatch.
- MAI-Voice-2.1 voice cloning is gated; how Sume TTS picks a voice
MAI-Voice-2.1 clones a voice from a 5 to 60 second clip but needs gated access. Sume TTS selects voices by avatar or voice id and has no reference-audio field.
- MAI-Voice-2.1 emotion control vs Sume TTS emotion, speed and volume
MAI-Voice-2.1 lists emotion control. Sume TTS takes generation_config: emotion (1-64 chars), speed 0.6-1.5, volume 0.5-2. A three-take test costs 3 cents.
- MAI-Voice-2.1 zero-shot voice prompting vs a Sume voice id or avatar
MAI-Voice-2.1 lists zero-shot voice prompting. Sume TTS takes a voice id or an avatar's voice, not a sample clip. What each request can and cannot express.
Written by Sume