MAI-Voice-2.1: 23 languages and voice matching; Sume takes a voice id
Microsoft says MAI-Voice-2.1 covers 23 languages and matches a voice from a short clip. Sume TTS takes a ready voice id or avatar, not reference audio.

Microsoft's MAI-Voice page lists 23 languages for MAI-Voice-2.1 and describes "instant voice matching" that captures a voice from a short reference clip with no fine-tuning. It does not state how long the clip must be. Sume TTS 1.0 has no reference-audio field: you send a transcript plus a ready voice id or an avatar, and the voice must be one Sume already holds.
Microsoft's claims are from its MAI-Voice page, read 2026-10-01; Sume's from the OpenAPI schema behind the API reference.
What does a Sume voice selector accept?
Either avatar_id / avatar_handle (Sume resolves that avatar's voice, which works when its voice status is ready) or voice.id. A voice id must be a TTS voice UUID or a Voices library id starting voi_. Anything else fails with 400 invalid_voice_id before a job is queued or credits are reserved.
How do the two models compare on input?
| Item | MAI-Voice-2.1 | Sume TTS 1.0 |
|---|---|---|
| Languages | 23 listed | Language field per request, 2 to 16 characters |
| Voice from a clip | Instant voice matching, length not stated | Not in the TTS request |
| Voice selector | Not described on this page | Voice UUID, voi_ id or avatar reference |
What does a request with a stored voice look like?
Pass the voice and the language the text is written in.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Bonjour et bienvenue.",
"voice": { "id": "voi_0123456789abcdef0123456789abcdef" },
"language": "fr"
}'Can I clone a voice from a clip on Sume?
Not through the TTS request. Voice cloning API covers what exists today. If you hold a consented voice elsewhere, check the rules of both vendors before reusing it.
Sources
Related posts
More in Developers
- MAI-Voice-2.1-Flash lists 45 ms; Sume TTS is an async job
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms and MAI-Voice-2.1 at about 550 ms. Sume TTS 1.0 is non-streaming: a job with a poll URL or webhook.
- Make AI Agent fallback connection retries once: key Sume calls
Make now retries an AI Agent run once on a fallback connection. Keep the Sume Idempotency-Key out of the model's hands so a retried run cannot bill twice.
- Make a voice louder than the music: gain_db and duck_db ranges
In Timeline 1.0, raise the voice with audio.gain_db (-60 to 12) and lower the music bed with soundtrack.duck_db (0 to 20). Ranges and refusal codes.
- Mastra MCP 2.0 is 2026-07-28 only: which line for Sume?
@mastra/mcp 2.0.0 drops the legacy handshake and speaks MCP 2026-07-28 only. The Sume server defaults to 2025-11-25, so pin the version you test.
Written by Sume