MAI-Voice-2.1: 14 serving regions, public preview, no SLA

MAI-Voice-2.1 is in public preview with no SLA and 14 serving regions. A pre-launch checklist, plus how an async TTS job changes the risk.

5 min readSume
All posts

Is MAI-Voice-2.1 ready for production? Not by Microsoft's own wording: the Microsoft Learn page says the feature is in public preview, provided without a service-level agreement, and not recommended for production workloads. Azure serves it from 14 regions, and the page says you can access it globally.

That is a launch-planning fact, not a quality verdict. Below is what the page says, and a checklist for deciding whether a preview voice belongs in your pipeline yet.

What the Learn page states

The page covers MAI-Voice-2.1 and MAI-Voice-2.1-Flash together. Both use the same Azure Speech API and SDK as other Azure neural and HD voices, so a MAI voice is chosen by name inside an SSML voice element rather than through a new endpoint.

Microsoft's announcement lists $22 per 1M characters for MAI-Voice-2.1 and $15 per 1M characters for Flash on the Microsoft AI announcement. The Learn page itself sends you to the Azure Speech pricing page rather than printing a price.

MAI-Voice-2.1 serving regions listed on Microsoft Learn (read 2026-10-03)
AreaRegion identifiers
United States and Canadaeastus, eastus2, westus, westus2, westus3, canadacentral
Europefrancecentral, westeurope, northeurope, swedencentral
Asia Pacificeastasia, southeastasia, centralindia, japaneast

A pre-launch checklist

Treat each line as a yes or no before real traffic hits a preview voice.

  • Contract: does your customer agreement require an SLA for the voice step? A preview has none, so the answer decides everything else.
  • Region: the page says you can access the models globally and that Azure routes requests to the 14 listed regions, but it still asks for a Foundry resource for Speech plus its key and region. Confirm data-residency needs against the page before you pick a resource region.
  • Fallback: can the pipeline switch to another voice if the preview voice errors or changes? Keep the fallback voice id in config, not in code.
  • Cloning: instant voice cloning is gated behind a Limited Access review, so a launch that depends on a brand voice has an approval step ahead of it.
  • Licensing: the page says Microsoft holds full licensing rights for commercial use. Read it with your legal team before you resell the audio.

Where an async job changes the risk

A real-time agent notices a failing voice within one turn. A batch narration job does not need to: you can regenerate a failed clip later. That makes batch work the safer first use of a preview model.

Sume's own TTS surface is job-based. A submit returns a job, and you read the result by polling or by webhook, as described in Sume jobs and results. Synchronous mode waits at most 30 seconds, and 30 seconds is a wait budget rather than a job duration. Sume TTS 1.0 is listed at $0.0475 per 1,000 characters in the Sume catalog, a figure you can read live from GET /v1/catalog (see the API reference).

None of that makes Sume a substitute for MAI-Voice-2.1; it is a different layer. It does mean a narration pipeline can keep its audio step behind one job interface, so swapping engines does not rewrite the rest. See the Flash latency comparison for the real-time side.

What to re-check after launch day

Previews change. The Learn page notes that Microsoft adds locales and managed voices as they become available, so re-read the page before each release instead of caching the voice list. Record the date you last read it next to your voice ids.

Sources

Related posts

More in Models

All Models posts

Written by Sume