MAI-Voice-2.1 on OpenRouter and Vercel: is it in Sume's TTS Router?
Microsoft lists MAI voices on Foundry, OpenRouter and Vercel. Sume's TTS Router lists only Cartesia Sonic ids today. How to check the live catalog.

No. Sume's TTS Router does not list MAI-Voice-2.1 or MAI-Voice-2.1-Flash. Its documented catalog is Cartesia Sonic only: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. Microsoft's launch post says the MAI voice models are available through Microsoft Foundry, the MAI Playground, Vercel and OpenRouter, with LiveKit coming soon (read 2026-10-03). If you want MAI voices today, call them through one of those, not through Sume.
That is not a permanent statement about Sume. The router docs say later vendors are additive catalog rows rather than a new surface. The way to know what is current is to read the catalog, which is free and takes one request.
Where Microsoft says the voices are
Prices on the model page are $22 per 1M characters for MAI-Voice-2.1 and $15 per 1M for Flash, with about 550 ms and about 45 ms of model inference. Those prices are Microsoft's; a gateway may add its own margin, so read the gateway's rate card before you budget.
| Route | What Microsoft says | Good for |
|---|---|---|
| Microsoft Foundry (Azure Speech) | Production access to MAI-Voice-2.1 and Flash | Azure-governed workloads |
| MAI Playground | Try the models in a browser | Listening tests |
| Copilot Audio Expressions | Listed on the model page as an access point | Consumer use |
| Vercel and OpenRouter | Listed in the launch post | Gateways you may already use |
| LiveKit | Coming soon | Voice agent rooms |
What Sume's router does and does not do
The Sume API reference describes POST /v1/tts-router/generate as a pass-through surface: you choose a catalog model, supply exactly one of transcript or transcript_source, pick a voice by avatar or voice id, and get an async job. The catalog rows carry list price per character times 1.25. Streaming TTS, other families and a model picker in the app are listed as out of scope for the first ship.
So the router is a fit when you want one job API for TTS, captions, music and timeline work with one bill, and you are content with Cartesia Sonic voices. It is not a fit when your requirement is a specific other vendor's voice.
Check the live catalog
Do not rely on a blog post for this, including this one. The catalog endpoint lists what is available right now. Run it, and look at the id of every row.
import json
import os
import urllib.request
req = urllib.request.Request(
"https://api.sume.com/v1/tts-router/models",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
)
with urllib.request.urlopen(req, timeout=30) as resp:
catalog = json.load(resp)
rows = catalog["data"] if isinstance(catalog, dict) else catalog
for row in rows:
print(row["id"])
wanted = [r for r in rows if "mai" in r["id"].lower()]
print("MAI rows:", len(wanted))
Questions to ask any TTS gateway
Whether you use a gateway for MAI voices or Sume's router for Sonic, the same five questions sort out the real differences. What is the margin over the vendor's list price, and is it published per model? Is the model id stable, or does it silently move to a newer version? Is the output a streamed byte stream or a durable file with an id you can fetch again? What are the commercial-use and disclosure terms for the voices? And what happens on failure: is a failed generation billed?
Sume's answers, from its docs: the catalog publishes list price per character and the 1.25 multiplier; sonic-latest is an alias for the current stable release while sonic-3.6 is the id to pin; the output is a job with a result you can read again; and failed or rejected requests are handled through the job statuses described in Jobs and results. Ask the same of any gateway and write the answers down next to the date. They go stale quickly in a month with new models every week.
A sensible split
Many teams will not pick one vendor. Use a gateway or Foundry for live, conversational speech where 45 to 150 ms matters. Use Sume where the output is a finished file you attach to a render: a voiceover for a timeline, a narration to caption, a track to join with others. The two meet at a file. If your MAI audio is a file, it can be imported to the Sume media host and used in timeline audio like any other audio, per that page's import-first rule. Check the license terms of the route you used for commercial and disclosure requirements before you publish.
Sources
Related posts
More in Models
- MiniMax H3 2K and 4K upscale on Sume: H3 accepts them, H3 Max does not
minimax-h3 will price a 2K or 4K request even though its resolution list shows only 480p and 768p. minimax-h3-max rejects both. Costs for 5 to 15 seconds.
- MiniMax H3 or H3 Max on Sume: which id for a 5 to 15 second clip?
Choose minimax-h3 for a 480p or 768p draft at $0.075 a second, minimax-h3-max when you need 1080p. Same 5-15 s range, same references; Max costs more at 768p.
- MiniMax H3 reference-to-video: five images free, then $0.08 each
On the minimax-h3 id, reference-to-video adds a list fee for every reference image after the fifth. minimax-h3-max does not. Worked 10-second totals.
- MiniMax-H3 Turbo LoRA: 4 to 8 steps, Apache 2.0, base licence
The MiniMax-H3 Turbo LoRA cuts sampling to 4 to 8 steps and lists Apache 2.0, but the 33B base keeps its own terms. What it needs and the hosted route.
Written by Sume