MAI-Transcribe-1.5 at $0.36 an hour vs Transcribe-2 and Streaming

Microsoft's own table lists MAI-Transcribe-1.5 at $0.36 per hour and 43 languages. Compare it with Transcribe-2 at $0.10, Streaming at $0.54 and Sume STT.

5 min readSume
All posts

On Microsoft's MAI-Transcribe-2 page, the version table lists MAI-Transcribe-1.5 at $0.36 per hour of audio and 43 languages. MAI-Transcribe-2 is listed at an introductory $0.10 per hour and 60 languages, and MAI-Transcribe-2-Streaming at $0.54 per hour. Sume STT 1.0 is $0.01 per audio minute, which is $0.60 per hour (read 2026-10-07).

The three rows side by side

The Version Comparison block on microsoft.ai has one card per model. The rows below copy it. The two prices marked introductory are the ones to watch; the Azure pricing page mentions a promotional offer for MAI Transcribe-2 until 12/31/2026, and I could not read a price there.

MAI-Transcribe versions as listed by Microsoft (read 2026-10-07)
1.522-Streaming
Artificial Analysis accuracy rank#3#2#1
Languages436060
Real-time streamingNoNoYes
Word-level timestampsNoYesNo
DiarizationNoYesNo
Keyword biasingNoYesNo
1 hour of audio finishes in20 sec10 secn/a (about 120 ms to first partial)
Price per hour$0.36$0.10 (introductory)$0.54 (introductory)

What 100 hours costs

Multiply the hourly price by the hours. For 100 hours of audio: 1.5 is 100 x $0.36 = $36.00, Transcribe-2 is 100 x $0.10 = $10.00, Streaming is 100 x $0.54 = $54.00, and Sume STT is 100 x 60 x $0.01 = $60.00. Transcribe-2 is the cheapest row in the table and also the only batch row with timestamps, diarization and biasing, so 1.5 is the odd one out: it is 3.6 times the price of 2 and has fewer languages.

  • Use Transcribe-2 for recorded files if you need timestamps or speaker labels and can live on the introductory price.
  • Use Streaming only when text must appear while someone is still talking.
  • Treat 1.5 as a legacy row: check that it is still deployable in your region before you plan on it.

What Sume STT covers

Sume STT 1.0 takes a public HTTPS audio_url and returns the text, a language code and word timings; the timings are always on. You can add sentence segmentation, and give a language_code hint or leave it empty for auto-detect. It has no keyword-biasing field and no style switch between clean and verbatim, and the diarization setting is fixed server-side with no speaker field in the request.

A job is billed at $0.01 per audio minute, capped at 10 minutes, so a longer recording is split first. That makes Sume the more expensive row on a pure per-hour basis, but it comes with the rest of the media chain in the same API: audio detach to pull the track from a video, and captions or a timeline afterwards.

Which to pick

If you only compare price, MAI-Transcribe-2 wins until the offer ends. If you need the transcript inside a media workflow that also cuts, captions and renders, the extra $0.50 per hour of Sume STT is the price of not wiring a second vendor. Re-read the Microsoft table before you commit: it is dated 2026-10-07 here and carries the words introductory twice.

One more check before you budget: the accuracy ranks are Microsoft's, taken from the Artificial Analysis leaderboard as quoted on its own page. I did not rerun any benchmark. A rank tells you little about your audio, so transcribe five of your own files with each candidate and count the wrong words per 1,000 before you choose. On Sume that five-file test is 5 cents if each clip is a minute long (5 x 1 minute x $0.01).

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume