MAI-Voice-2.1-Flash: Voice Live or the Speech SDK with SSML?
Microsoft's docs show two ways to call MAI-Voice-2.1-Flash: the Azure Speech SDK with SSML, and Voice Live. Which fits an agent, which fits a batch.

Microsoft's documentation says you can integrate MAI-Voice-2.1-Flash with the Azure Speech SDK through SSML, and also through Voice Live. Use the SDK or REST with SSML for batches of lines and files, and Voice Live when the speech is part of a real-time conversation.
What the docs show
The Microsoft Learn page, updated October 1, 2026, marks the feature as public preview without an SLA and not recommended for production workloads. Key points:
- Both models use the same Azure Speech API and SDK as other neural voices, with a MAI voice name in the SSML voice element.
- The voice name carries a model suffix, for example en-US-Harper:MAI-Voice-2.1-Flash.
- Style control uses mstts:express-as with a style attribute.
- Instant voice cloning is gated, with a recommended 5 to 60 second reference clip.
- Pricing is on the Azure Speech pricing page; Microsoft's announcement lists $15 per million characters for Flash and $22 for MAI-Voice-2.1.
Choosing the route
Match the route to the shape of the work.
| Work | Route | Reason |
|---|---|---|
| 50 ad lines to MP3 files | Speech SDK or REST with SSML | Request in, file out |
| A voice assistant on a phone line | Voice Live | Streaming, turn-taking |
| Hundreds of lines overnight | SDK in a loop | Predictable and cheap |
| One-off test in a browser | Foundry playground | No code |
The same split exists on Sume
Sume's TTS is the file-oriented route. You submit tts_create, get a job id, and read the audio from the job result. At $0.0475 per 1,000 characters it costs more per character than Microsoft's $15 per million for Flash, so choose on what you need in the file and in the surrounding chain.
Preview caution
Because Microsoft marks the feature as preview, do not ship an unreviewed pipeline on it. Pin the voice name with the model suffix, keep the approved audio files, and re-listen after any update.
Sources
Related posts
More in Developers
- make -j4 as a resumable Wan 3.0 batch runner with an Idempotency-Key
Twelve Wan 3.0 shots as twelve make targets. Re-run after a crash and only missing outputs submit, with an Idempotency-Key per shot. Tested on a stub server.
- Mandarin Chinese speech to text API: Sume STT language_code zh
Transcribe Mandarin audio with Sume STT using language_code zh, then check the result and timings. $0.01 per audio minute and no accuracy claim without a test.
- mask_url on ChatGPT Image 2: not listed, only the 2.5 rows take it
Sume lists mask_url only on ChatGPT Image 2.5 Flare and Sunburst. Sending it to ChatGPT Image 2 is rejected. What to do for a masked edit on the older row.
- MCP outputSchema vs Sume output_schema: who sets the contract
In MCP the server declares a tool's outputSchema. In a Sume Agent Completion you send output_schema per run, and the result can still come back degraded.
Written by Sume