MAI-Transcribe-1 deprecated Aug 20: what to switch to, and Sume STT

Microsoft's page says MAI-Transcribe-1 was deprecated on August 20, 2026. What replaces it, what is still preview, and what a Sume STT job needs instead.

4 min readSume
All posts

What replaces MAI-Transcribe-1

Move to MAI-Transcribe-2 or MAI-Transcribe-1.5. Microsoft's MAI-Transcribe page (read 2026-10-04) lists those two models and states that MAI-Transcribe-1 was deprecated on August 20, 2026. The same page marks the service as public preview with no SLA, so a model change can happen again.

If you call the model by name, update that string and re-test your output. Look at transcribe style, timestamps and locales, because those are options on the request.

What the Microsoft page says

The facts below come from the page itself, not from us.

MAI-Transcribe facts from Microsoft Learn (read 2026-10-04)
ItemWhat the page says
Models listedMAI-Transcribe-2 and MAI-Transcribe-1.5
MAI-Transcribe-1Deprecated August 20, 2026
StatusPublic preview, no SLA
Input formatsWAV, MP3 or FLAC
Endpoint versionapi-version 2025-10-15
OptionsDiarization, word timestamps, phrase list, verbatim or clean style

The Sume side of the same problem

Sume STT has no model field to migrate. A job is POST /v1/stt-1.0/transcribe with the public model sume/stt-1.0, so a vendor retiring a model name does not change your request. You send audio_url, an optional language_code and an optional duration_seconds of 1 to 600.

Word timings always come back in words[] as {word, start, end}, and there is no flag to switch them on. Diarization and event tagging are fixed on the server, so you cannot choose a transcribe style or a phrase list. If you need those controls, Microsoft's service has them; if you want the same request shape year after year, a fixed route is simpler.

A migration checklist

Moving off a deprecated speech model is mostly a testing exercise. The request shape changes less than the output does. Different models can punctuate differently, break words at different points and return different timings, so anything downstream that depends on those details needs a re-test.

  • Collect five real files: clean audio, noisy audio, a phone call, a multi-speaker clip and a code-switched clip.
  • Transcribe them with the new model and diff against your current output, not against a reference you made up.
  • Check that caption cues, search indexes and any word-level alignment still line up.
  • Keep the old output for a week in case a customer disputes a transcript.
  • Note the endpoint version in your config, since the page shows api-version 2025-10-15 today.

Price check

Sume STT is $0.01 per audio minute, so one hour is $0.60. Microsoft's introductory price for its streaming model is on a different page and for a different product; compare it to your own usage rather than to a single number. For an hour of audio a month the difference is cents; for thousands of hours, the per-hour rate and the preview risk both matter. Price your real volume, then decide, and remember that a lower rate with no SLA is a different product from a fixed route with job ids you can poll, cancel and audit.

  • Pin nothing in your code that depends on a deprecated name.
  • Run a sample file through the new model before you cut over.
  • Treat preview services as changeable and keep a fallback route.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume