MAI-Voice-2.1 Python quickstart: Speech SDK and SSML to MP3

Microsoft's MAI-Voice-2.1 Python sample uses the Azure Speech SDK, SSML and a 24 kHz MP3 output. The code, the settings and what to change.

5 min readSume
All posts

How do you call MAI-Voice-2.1 from Python? Install the Azure Speech SDK, set SPEECH_KEY and SPEECH_REGION, and send SSML naming a MAI voice such as en-US-Harper:MAI-Voice-2.1. The Microsoft Learn page gives this sample, and it writes a 24 kHz mono MP3 at 160 kbps.

The same page has tabs for the Foundry portal, REST, C#, JavaScript and Java, so the model is not tied to one language.

The sample

This is Microsoft's own example, with the voice set to the full model. It needs pip install azure-cognitiveservices-speech, a Foundry resource for Speech and its region.

import os

import azure.cognitiveservices.speech as speechsdk

speech_config = speechsdk.SpeechConfig(
    subscription=os.environ["SPEECH_KEY"],
    region=os.environ["SPEECH_REGION"],
)
speech_config.set_speech_synthesis_output_format(
    speechsdk.SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3
)
audio_config = speechsdk.audio.AudioOutputConfig(filename="output.mp3")
synthesizer = speechsdk.SpeechSynthesizer(
    speech_config=speech_config,
    audio_config=audio_config,
)

ssml = """
<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xml:lang="en-US">
  <voice name="en-US-Harper:MAI-Voice-2.1">
    Hello, this is a sample from MAI Voice.
  </voice>
</speak>
"""

result = synthesizer.speak_ssml_async(ssml).get()
if result.reason != speechsdk.ResultReason.SynthesizingAudioCompleted:
    raise RuntimeError(f"Speech synthesis failed: {result.reason}")

What each setting does

Four lines decide the outcome, and each is easy to change.

Settings in the Microsoft MAI-Voice Python sample (read 2026-10-03)
SettingValue in the sampleWhat to change
SPEECH_KEYYour Speech resource key, from an environment variableKeep it in the environment, never in code
SPEECH_REGIONYour resource regionYour resource region; the page lists 14 regions that serve the models and says access is global
Output formatAudio24Khz160KBitRateMonoMp3Pick another SpeechSynthesisOutputFormat for WAV or other rates
Voice nameen-US-Harper:MAI-Voice-2.1Use the -Flash suffix for the low-latency model

Switching to Flash or another voice

Change only the SSML voice name. The page says the same suffix applies to any supported prebuilt voice, so en-US-Harper:MAI-Voice-2.1-Flash selects Flash on the same voice. For a style, wrap the text in mstts:express-as with a style that the voice lists; the voice and style lists show how those differ by voice.

The sample checks result.reason and raises on failure. Keep that check, since a preview service can reject a request for reasons that are not in your code.

Prerequisites and the REST alternative

You need an Azure subscription and a Microsoft Foundry resource for Speech, plus its key and region. If you would rather skip the SDK, the same page shows a REST call: an SSML POST to the cognitiveservices/v1 endpoint of your region, with an X-Microsoft-OutputFormat header such as audio-24khz-160kbitrate-mono-mp3 and your key in the Ocp-Apim-Subscription-Key header. The response body is the audio, so you save it straight to a file.

For a style, the SSML root must also declare the mstts namespace, as Microsoft's own style example does. The basic sample above leaves it out because it uses no style.

Batch work and job interfaces

This call is synchronous from your script's point of view: the audio is written when the call returns. For long scripts, split the text and run several calls, then join the files. If you prefer to hand the work to a job interface, Sume's TTS endpoints return a job to poll or receive by webhook, with a synchronous wait capped at 30 seconds, per Sume jobs and results.

Whichever path you take, keep the voice name, model suffix and output format in config, so a change of model is one edit.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume