MAI-Voice-2.1 on Foundry and Vercel; Sume's TTS Router is Sonic only
Microsoft lists Foundry, MAI Playground, Vercel and OpenRouter for MAI-Voice-2.1. Sume's TTS Router lists only Cartesia Sonic ids, so check the catalog.

Microsoft's October 1 post says MAI-Voice-2.1 and its Flash variant are available in Microsoft Foundry, the MAI Playground, Vercel and OpenRouter, with LiveKit coming soon. Sume does not offer MAI-Voice-2.1. The TTS Router is Cartesia Sonic only, and the five ids are sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview.
How to check for yourself
Do not look for an MAI model id on a Sume route. An unknown model fails with 400 model_not_found and the catalog URL. The way to know what is on offer today is GET /v1/tts-router/models. The router docs say later vendors would be added as new catalog rows, so read the list and not this page for the current set.
What to do with that
If you want to try an MAI voice, use one of the places Microsoft names. If you want a voice file that comes back as a Sume job, with a Sume job URL, webhook delivery and one idempotency key, use TTS 1.0. The two are not exclusive. You can compare a line from each by ear before you commit a project to either.
The two offers, read 2026-10-06:
| Item | MAI-Voice-2.1 family (vendor) | Sume TTS Router (docs) |
|---|---|---|
| Places | Foundry, MAI Playground, Vercel, OpenRouter | One Sume API |
| Models | MAI-Voice-2.1, MAI-Voice-2.1-Flash | Five Sonic ids |
| Flash latency claim | 150 ms end to end for 45 s of audio | Async job, no streaming |
| Selection | Vendor model name | Required model id from the catalog |
A fair side-by-side
Pick three lines from your real script: a short hook, a number-heavy line and a long sentence. Generate each in both services, with the same language, and listen on the speakers your audience will use. Record price and wait time in your own sheet.
Do not choose from a demo page. Vendor demos use their best lines. Your script has names, units and odd words, and those are where voices differ.
Also write down what you need beyond the voice. A Sume job gives you a status URL, a result URL, signed webhook delivery and an idempotency key on every write, and its audio can go straight into a timeline render. A vendor playground gives you a good listen. Decide which of those matters for your project, and test that part too.
Sume's TTS is an async job with a poll or webhook, so a vendor's end-to-end latency figure is not a number you can expect from it. Compare on the lines you will really ship, and measure from submit to a finished file.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 or Sume TTS? Three questions that decide it
Live speech, extra outputs or lowest price per character? A short guide to MAI-Voice-2.1 and Flash against the Sume TTS Router, with 10M-character math.
- Shortest AI video clip by model: MiniMax H3 4 s, Sume row 5 s
MiniMax lists 4 to 15 seconds for H3. Sume's row starts at 5. Here are the minimum durations in the catalog and how to trim a longer clip.
- MiniMax H3 open weights vs a hosted lip-sync API: what you take on
MiniMax released H3 open weights on 2026-08-03. Self-hosting is not the same as a hosted still-plus-audio lip-sync route. What each choice makes you own.
- Notion Agent SDK beta vs Sume Agent Completions: run agents from code
Notion's SDK continues agent chats and streams results. Sume's Agent Completions are one-shot 202 runs you poll or get by webhook. See which fits your backend.
Written by Sume