MAI-Voice-2.1 Python quickstart: Speech SDK and SSML to MP3
Microsoft's MAI-Voice-2.1 Python sample uses the Azure Speech SDK, SSML and a 24 kHz MP3 output. The code, the settings and what to change.

How do you call MAI-Voice-2.1 from Python? Install the Azure Speech SDK, set SPEECH_KEY and SPEECH_REGION, and send SSML naming a MAI voice such as en-US-Harper:MAI-Voice-2.1. The Microsoft Learn page gives this sample, and it writes a 24 kHz mono MP3 at 160 kbps.
The same page has tabs for the Foundry portal, REST, C#, JavaScript and Java, so the model is not tied to one language.
The sample
This is Microsoft's own example, with the voice set to the full model. It needs pip install azure-cognitiveservices-speech, a Foundry resource for Speech and its region.
import os
import azure.cognitiveservices.speech as speechsdk
speech_config = speechsdk.SpeechConfig(
subscription=os.environ["SPEECH_KEY"],
region=os.environ["SPEECH_REGION"],
)
speech_config.set_speech_synthesis_output_format(
speechsdk.SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3
)
audio_config = speechsdk.audio.AudioOutputConfig(filename="output.mp3")
synthesizer = speechsdk.SpeechSynthesizer(
speech_config=speech_config,
audio_config=audio_config,
)
ssml = """
<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xml:lang="en-US">
<voice name="en-US-Harper:MAI-Voice-2.1">
Hello, this is a sample from MAI Voice.
</voice>
</speak>
"""
result = synthesizer.speak_ssml_async(ssml).get()
if result.reason != speechsdk.ResultReason.SynthesizingAudioCompleted:
raise RuntimeError(f"Speech synthesis failed: {result.reason}")What each setting does
Four lines decide the outcome, and each is easy to change.
| Setting | Value in the sample | What to change |
|---|---|---|
| SPEECH_KEY | Your Speech resource key, from an environment variable | Keep it in the environment, never in code |
| SPEECH_REGION | Your resource region | Your resource region; the page lists 14 regions that serve the models and says access is global |
| Output format | Audio24Khz160KBitRateMonoMp3 | Pick another SpeechSynthesisOutputFormat for WAV or other rates |
| Voice name | en-US-Harper:MAI-Voice-2.1 | Use the -Flash suffix for the low-latency model |
Switching to Flash or another voice
Change only the SSML voice name. The page says the same suffix applies to any supported prebuilt voice, so en-US-Harper:MAI-Voice-2.1-Flash selects Flash on the same voice. For a style, wrap the text in mstts:express-as with a style that the voice lists; the voice and style lists show how those differ by voice.
The sample checks result.reason and raises on failure. Keep that check, since a preview service can reject a request for reasons that are not in your code.
Prerequisites and the REST alternative
You need an Azure subscription and a Microsoft Foundry resource for Speech, plus its key and region. If you would rather skip the SDK, the same page shows a REST call: an SSML POST to the cognitiveservices/v1 endpoint of your region, with an X-Microsoft-OutputFormat header such as audio-24khz-160kbitrate-mono-mp3 and your key in the Ocp-Apim-Subscription-Key header. The response body is the audio, so you save it straight to a file.
For a style, the SSML root must also declare the mstts namespace, as Microsoft's own style example does. The basic sample above leaves it out because it uses no style.
Batch work and job interfaces
This call is synchronous from your script's point of view: the audio is written when the call returns. For long scripts, split the text and run several calls, then join the files. If you prefer to hand the work to a job interface, Sume's TTS endpoints return a job to poll or receive by webhook, with a synchronous wait capped at 30 seconds, per Sume jobs and results.
Whichever path you take, keep the voice name, model suffix and output format in config, so a change of model is one edit.
Sources
Related posts
More in Developers
- Make an AI voice read numbers, prices and years correctly: a test list
Test how a Sume TTS voice reads prices, years, percentages and phone numbers: eight lines in one job, sentence slices to audition, and a respell fallback.
- Match loudness between voiceover takes: test sentence plus volume
Two takes recorded on different days sound different in level. Measure one fixed test sentence, then set Sume TTS generation_config.volume (0.5 to 2.0).
- MAX_MCP_OUTPUT_TOKENS 25,000: size a Sume jobs_wait wave read
Claude Code caps MCP output at 25,000 tokens and saves larger results to a file. How that meets Sume's jobs_wait include_results and a batch jobs_result read.
- MCP 401 challenge: the WWW-Authenticate header Sume returns
What a client sees when it calls Sume's MCP endpoint with no token: the WWW-Authenticate challenge, its resource_metadata URL and scope, and the spec.
Written by Sume