Where can I use MAI-Voice-2.1 today, and what Sume offers instead
Microsoft lists its Playground, Copilot Audio Expressions and Foundry; the launch post adds OpenRouter and Vercel. What Sume's text to speech gives you instead.

Where can you actually use MAI-Voice-2.1 right now? Microsoft's model page (read 2026-10-04) names three places: the MAI Playground, Copilot Audio Expressions, and Microsoft Foundry through Azure Speech. The launch post (read 2026-10-04) also lists OpenRouter and Vercel, and says LiveKit and Azure Voice Live are upcoming.
Sume does not list MAI-Voice-2.1. If you want that voice, use one of Microsoft's routes. If you want text to speech that feeds a captioned video, Sume's own routes are below.
Access, side by side
Everything in the Microsoft rows is from the two Microsoft pages named above.
| Route | What it is | Needs |
|---|---|---|
| MAI Playground | Microsoft's try-it surface for MAI-Voice-2.1 | A browser |
| Copilot Audio Expressions | Microsoft's Copilot feature that uses the model | A Copilot account |
| Microsoft Foundry (Azure Speech) | API access through Azure | An Azure account |
| OpenRouter, Vercel | Listed in Microsoft's launch post | An account with that gateway |
| Sume TTS 1.0 | POST /v1/tts-1.0/generate, async job, hosted audio | A Sume API key |
| Sume TTS Router | POST /v1/tts-router/generate with an explicit model from the catalog (Sonic ids) | A Sume API key |
What Sume gives you
A TTS request takes up to 20,000 characters, a language for any non-English text, and a voice chosen by voice.id or by avatar_id / avatar_handle. Output defaults to mp3 at 44.1 kHz and 128 kbps; ask for wav when the audio will be joined or lip-synced. Optional timestamps.words returns word timings, and segmentation can cut gapless sentence slices. Audio longer than 1200 seconds fails with tts_duration_exceeded and no credit is captured. These limits are in the OpenAPI contract behind the API reference.
Hosted MCP exposes the same family as tools, tts_create and stt_create, listed in MCP tools and gates, so an assistant that can call Sume can voice a script without you writing code.
Which one to pick
Choose Microsoft's routes when the specific model, its 23 languages or its Copilot integration is the point. Choose Sume when the audio is one input to a larger job: word timings for captions, sentence slices for scenes, a spine for a timeline render, or a voice tied to an avatar.
Check the live model list before you assume a new launch is available: GET /v1/tts-router/models is the source of truth, and an unknown id fails with 400 model_not_found.
Sources
Related posts
More in Models
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
- Music generation API: the Sume Music Router with Lyria 3.5
Sume's Music Router turns a text prompt into a track via POST /v1/music-router/generate. sume/music-auto picks the engine, Lyria 3.5 today.
Written by Sume