MAI-Transcribe-1 deprecated Aug 20: what to switch to, and Sume STT
Microsoft's page says MAI-Transcribe-1 was deprecated on August 20, 2026. What replaces it, what is still preview, and what a Sume STT job needs instead.

What replaces MAI-Transcribe-1
Move to MAI-Transcribe-2 or MAI-Transcribe-1.5. Microsoft's MAI-Transcribe page (read 2026-10-04) lists those two models and states that MAI-Transcribe-1 was deprecated on August 20, 2026. The same page marks the service as public preview with no SLA, so a model change can happen again.
If you call the model by name, update that string and re-test your output. Look at transcribe style, timestamps and locales, because those are options on the request.
What the Microsoft page says
The facts below come from the page itself, not from us.
| Item | What the page says |
|---|---|
| Models listed | MAI-Transcribe-2 and MAI-Transcribe-1.5 |
| MAI-Transcribe-1 | Deprecated August 20, 2026 |
| Status | Public preview, no SLA |
| Input formats | WAV, MP3 or FLAC |
| Endpoint version | api-version 2025-10-15 |
| Options | Diarization, word timestamps, phrase list, verbatim or clean style |
The Sume side of the same problem
Sume STT has no model field to migrate. A job is POST /v1/stt-1.0/transcribe with the public model sume/stt-1.0, so a vendor retiring a model name does not change your request. You send audio_url, an optional language_code and an optional duration_seconds of 1 to 600.
Word timings always come back in words[] as {word, start, end}, and there is no flag to switch them on. Diarization and event tagging are fixed on the server, so you cannot choose a transcribe style or a phrase list. If you need those controls, Microsoft's service has them; if you want the same request shape year after year, a fixed route is simpler.
A migration checklist
Moving off a deprecated speech model is mostly a testing exercise. The request shape changes less than the output does. Different models can punctuate differently, break words at different points and return different timings, so anything downstream that depends on those details needs a re-test.
- Collect five real files: clean audio, noisy audio, a phone call, a multi-speaker clip and a code-switched clip.
- Transcribe them with the new model and diff against your current output, not against a reference you made up.
- Check that caption cues, search indexes and any word-level alignment still line up.
- Keep the old output for a week in case a customer disputes a transcript.
- Note the endpoint version in your config, since the page shows api-version 2025-10-15 today.
Price check
Sume STT is $0.01 per audio minute, so one hour is $0.60. Microsoft's introductory price for its streaming model is on a different page and for a different product; compare it to your own usage rather than to a single number. For an hour of audio a month the difference is cents; for thousands of hours, the per-hour rate and the preview risk both matter. Price your real volume, then decide, and remember that a lower rate with no SLA is a different product from a fixed route with job ids you can poll, cancel and audit.
- Pin nothing in your code that depends on a deprecated name.
- Run a sample file through the new model before you cut over.
- Treat preview services as changeable and keep a fallback route.
Sources
Related posts
More in Comparisons
- MAI-Transcribe-2 timestamps option vs Sume STT: words[] are always on
On Microsoft's MAI-Transcribe you ask for word timestamps in modelOptions. Sume STT returns words[] with start and end on every job: what that changes for you.
- MAI-Voice-2.1 lists Korean, Thai, Vietnamese: Sume's language field
Microsoft's MAI-Voice-2.1 page lists 23 languages including Korean, Thai and Vietnamese. What to check before you plan a non-English voiceover on Sume TTS.
- MAI-Voice-2.1 or Flash for explainer narration: latency barely matters
Microsoft lists about 550 ms vs 45 ms model inference. For narration rendered ahead of time, pick on quality and price, then see how Sume jobs return audio.
- MakeUGC API Starter $99 for 2,000 credits vs Sume avatar seconds
MakeUGC API Starter is $99 a month with 2,000 credits. The same $99 on Sume buys 13 Plus 30-second avatar clips at $7.35 each, or 17 on Standard at $5.52.
Written by Sume