Where MAI-Transcribe-2-Streaming runs today, and a file-based option
Microsoft lists Foundry, MAI Playground and Vercel for MAI-Transcribe-2-Streaming. Its $0.54/h intro rate beside Sume STT at $0.60/h, and when a file job fits.

On October 1, 2026 Microsoft AI announced MAI-Transcribe-2-Streaming. Its announcement says it is accessible through Microsoft Foundry, the MAI Playground and Vercel, with LiveKit and Azure Voice Live to follow. It lists 60 languages with automatic, continuous language detection, first partial transcripts in just over 100 ms, and $0.54 per hour of audio through the end of the year (read 2026-10-05).
What streaming means for you
A streaming model returns words while audio is still arriving. That is built for live captions and voice agents. If your audio is a finished file, such as a recorded interview, a call recording or a video you edited, you do not need partial results. You need an accurate transcript with timings and a predictable price.
Side by side
| Item | MAI-Transcribe-2-Streaming | Sume STT 1.0 |
|---|---|---|
| Shape | Streaming, partial transcripts | Async job on a public HTTPS audio_url |
| Price | $0.54 per hour, introductory, through year end | $0.01 per audio minute, $0.60 per hour |
| Languages | 60, automatic language detection | Optional language_code, omit to auto-detect |
| Word timings | Partials in just over 100 ms (a latency figure, not a timing format) | words[] on every result, sentence segments on request |
| Longest input | Live stream | 600 seconds per job |
How to read the price gap
At the listed rates, an hour is $0.54 against $0.60, a difference of 6 cents. The Microsoft rate is introductory and ends after 2026, so a budget that runs into 2027 should not assume it. For a 45-minute weekly recording the gap is under 5 cents.
When a file job is the better fit
Cut the video's audio with audio detach (1 cent per job), split it into pieces of at most 600 seconds, and submit each to Sume STT. Each job returns word timings you can use for captions, chapter markers or a searchable transcript. Sume's job envelope also gives you a job id, a status URL and an optional webhook, so a batch of 100 recordings is a loop with an idempotency key per file.
- Choose streaming when a person is waiting for partial text.
- Choose a file job when the audio already exists and you want timings, not latency.
Sources
Related posts
More in Comparisons
- Where to use Wan 3.0: wan.video, QwenCloud, or the wan-3.0 API on Sume
Which door to Wan 3.0 fits: Alibaba's channels named in its README, or the wan-3.0 id on Sume with an API key, MCP and Timeline for multi-shot work.
- Which MAI-Voice-2.1 model for ad voiceover: standard or Flash?
Microsoft positions MAI-Voice-2.1 for voice-over and audiobooks, Flash for live agents. For rendered ads pick standard; Sume is the async file path.
- Text change, region change or background swap: which Sume image route
Pick the Sume image edit route by the kind of change: Ideogram 4.5 for words, ChatGPT Image 2.5 with mask_url for a region, a reference edit for backgrounds.
- Which platforms auto-detect AI media: YouTube, Meta, Pinterest, TikTok
YouTube, Meta and Pinterest say they can label AI media from metadata or detection; TikTok auto-labels its own AI effects. None replaces your own disclosure.
Written by Sume