MAI-Voice-2.1-Flash: Voice Live or the Speech SDK with SSML?

Microsoft's docs show two ways to call MAI-Voice-2.1-Flash: the Azure Speech SDK with SSML, and Voice Live. Which fits an agent, which fits a batch.

5 min readSume
All posts

Microsoft's documentation says you can integrate MAI-Voice-2.1-Flash with the Azure Speech SDK through SSML, and also through Voice Live. Use the SDK or REST with SSML for batches of lines and files, and Voice Live when the speech is part of a real-time conversation.

What the docs show

The Microsoft Learn page, updated October 1, 2026, marks the feature as public preview without an SLA and not recommended for production workloads. Key points:

  • Both models use the same Azure Speech API and SDK as other neural voices, with a MAI voice name in the SSML voice element.
  • The voice name carries a model suffix, for example en-US-Harper:MAI-Voice-2.1-Flash.
  • Style control uses mstts:express-as with a style attribute.
  • Instant voice cloning is gated, with a recommended 5 to 60 second reference clip.
  • Pricing is on the Azure Speech pricing page; Microsoft's announcement lists $15 per million characters for Flash and $22 for MAI-Voice-2.1.

Choosing the route

Match the route to the shape of the work.

Integration routes (read 2026-10-07)
WorkRouteReason
50 ad lines to MP3 filesSpeech SDK or REST with SSMLRequest in, file out
A voice assistant on a phone lineVoice LiveStreaming, turn-taking
Hundreds of lines overnightSDK in a loopPredictable and cheap
One-off test in a browserFoundry playgroundNo code

The same split exists on Sume

Sume's TTS is the file-oriented route. You submit tts_create, get a job id, and read the audio from the job result. At $0.0475 per 1,000 characters it costs more per character than Microsoft's $15 per million for Flash, so choose on what you need in the file and in the surrounding chain.

Preview caution

Because Microsoft marks the feature as preview, do not ship an unreviewed pipeline on it. Pin the voice name with the model suffix, keep the approved audio files, and re-listen after any update.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume