Where MAI-Transcribe-2-Streaming runs today, and a file-based option

Microsoft lists Foundry, MAI Playground and Vercel for MAI-Transcribe-2-Streaming. Its $0.54/h intro rate beside Sume STT at $0.60/h, and when a file job fits.

4 min readSume
All posts

On October 1, 2026 Microsoft AI announced MAI-Transcribe-2-Streaming. Its announcement says it is accessible through Microsoft Foundry, the MAI Playground and Vercel, with LiveKit and Azure Voice Live to follow. It lists 60 languages with automatic, continuous language detection, first partial transcripts in just over 100 ms, and $0.54 per hour of audio through the end of the year (read 2026-10-05).

What streaming means for you

A streaming model returns words while audio is still arriving. That is built for live captions and voice agents. If your audio is a finished file, such as a recorded interview, a call recording or a video you edited, you do not need partial results. You need an accurate transcript with timings and a predictable price.

Side by side

MAI-Transcribe-2-Streaming (Microsoft, read 2026-10-05) vs Sume STT 1.0
ItemMAI-Transcribe-2-StreamingSume STT 1.0
ShapeStreaming, partial transcriptsAsync job on a public HTTPS audio_url
Price$0.54 per hour, introductory, through year end$0.01 per audio minute, $0.60 per hour
Languages60, automatic language detectionOptional language_code, omit to auto-detect
Word timingsPartials in just over 100 ms (a latency figure, not a timing format)words[] on every result, sentence segments on request
Longest inputLive stream600 seconds per job

How to read the price gap

At the listed rates, an hour is $0.54 against $0.60, a difference of 6 cents. The Microsoft rate is introductory and ends after 2026, so a budget that runs into 2027 should not assume it. For a 45-minute weekly recording the gap is under 5 cents.

When a file job is the better fit

Cut the video's audio with audio detach (1 cent per job), split it into pieces of at most 600 seconds, and submit each to Sume STT. Each job returns word timings you can use for captions, chapter markers or a searchable transcript. Sume's job envelope also gives you a job id, a status URL and an optional webhook, so a batch of 100 recordings is a loop with an idempotency key per file.

  • Choose streaming when a person is waiting for partial text.
  • Choose a file job when the audio already exists and you want timings, not latency.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume