Where can I use MAI-Voice-2.1 today, and what Sume offers instead

Microsoft lists its Playground, Copilot Audio Expressions and Foundry; the launch post adds OpenRouter and Vercel. What Sume's text to speech gives you instead.

5 min readSume
All posts

Where can you actually use MAI-Voice-2.1 right now? Microsoft's model page (read 2026-10-04) names three places: the MAI Playground, Copilot Audio Expressions, and Microsoft Foundry through Azure Speech. The launch post (read 2026-10-04) also lists OpenRouter and Vercel, and says LiveKit and Azure Voice Live are upcoming.

Sume does not list MAI-Voice-2.1. If you want that voice, use one of Microsoft's routes. If you want text to speech that feeds a captioned video, Sume's own routes are below.

Access, side by side

Everything in the Microsoft rows is from the two Microsoft pages named above.

Ways to reach a TTS model, read 2026-10-04
RouteWhat it isNeeds
MAI PlaygroundMicrosoft's try-it surface for MAI-Voice-2.1A browser
Copilot Audio ExpressionsMicrosoft's Copilot feature that uses the modelA Copilot account
Microsoft Foundry (Azure Speech)API access through AzureAn Azure account
OpenRouter, VercelListed in Microsoft's launch postAn account with that gateway
Sume TTS 1.0POST /v1/tts-1.0/generate, async job, hosted audioA Sume API key
Sume TTS RouterPOST /v1/tts-router/generate with an explicit model from the catalog (Sonic ids)A Sume API key

What Sume gives you

A TTS request takes up to 20,000 characters, a language for any non-English text, and a voice chosen by voice.id or by avatar_id / avatar_handle. Output defaults to mp3 at 44.1 kHz and 128 kbps; ask for wav when the audio will be joined or lip-synced. Optional timestamps.words returns word timings, and segmentation can cut gapless sentence slices. Audio longer than 1200 seconds fails with tts_duration_exceeded and no credit is captured. These limits are in the OpenAPI contract behind the API reference.

Hosted MCP exposes the same family as tools, tts_create and stt_create, listed in MCP tools and gates, so an assistant that can call Sume can voice a script without you writing code.

Which one to pick

Choose Microsoft's routes when the specific model, its 23 languages or its Copilot integration is the point. Choose Sume when the audio is one input to a larger job: word timings for captions, sentence slices for scenes, a spine for a timeline render, or a voice tied to an avatar.

Check the live model list before you assume a new launch is available: GET /v1/tts-router/models is the source of truth, and an unknown id fails with 400 model_not_found.

Sources

Related posts

More in Models

All Models posts

Written by Sume