MAI-Voice-2.1 or Sume TTS? Three questions that decide it
Live speech, extra outputs or lowest price per character? A short guide to MAI-Voice-2.1 and Flash against the Sume TTS Router, with 10M-character math.

Choose Microsoft's MAI-Voice-2.1 Flash if speech has to start inside a live turn, and MAI-Voice-2.1 if per-character price decides and you do not need anything else from the API. Choose Sume's TTS Router if you want a finished audio file that comes with word timings, sentence segments, a voice-language guard and signed webhooks on the same job, and if the same account also covers captions, lip sync and video. Three questions get you there. The facts below come from the October 2026 tracker (read 2026-10-06) and Sume's pricing page.
1. Does speech start inside a live turn?
The tracker lists MAI-Voice-2.1-Flash at a vendor-claimed 150 ms end to end. Sume TTS is an async job that returns a hosted file, and streaming TTS is a documented non-goal of the router. If a person is waiting for the first word, that settles it for the Microsoft product. If a person presses a button and waits a few seconds for a clip, it does not. Sync mode waits at most 30 seconds on one request and falls back to polling, so it fits a button, not a call.
2. How much of the bill is characters?
MAI-Voice-2.1 is listed at $22 per 1M characters and Flash at $15. Sume's router is $47.50 per 1M on every model, the provider list price with a 1.25 margin. At small volume the gap is cents. At 10M characters a month it is $220 against $475, or $150 for Flash. If the characters are the whole bill, that is the number to weigh. If you also pay for transcription, caption renders, a second vendor's webhooks and the glue between them, count those too.
3. What else do you need next to the audio?
A voiceover is rarely the final product. Sume's TTS job can return words[] with start and end times, and gapless sentence segments[] with a 70 ms default boundary lead. The completed job records the model, voice, language and output settings. A wrong-language voice stops with a 409 before it spends. The same account runs speech-to-text at $0.01 per audio minute, captions at $0.20 per job and Timeline audio at $0.01 per job. Whether that matters is a question about your pipeline, not about the voice.
| If you need | MAI-Voice-2.1 | MAI-Voice-2.1-Flash | Sume TTS Router |
|---|---|---|---|
| Speech inside a live turn | Not stated | Vendor claims 150 ms end to end | No streaming |
| Lowest characters bill, 10M a month | $220 | $150 | $475 |
| Word timings and sentence segments with the audio | Not stated in the tracker | Not stated in the tracker | Yes, one request |
| Languages | 23 languages, 26 locales | Not stated separately | Set language, guard checks the voice |
| Captions, lip sync and video on one key | Not in the tracker | Not in the tracker | Yes |
Use both, and check the numbers
Many teams end up with both: a streaming model for the live path and Sume for everything that ends as a file. Keep the boundary clean. Decide per surface, not per company, and put the choice behind one function so you can swap a voice without touching the caller.
One caution about the table: the Microsoft figures are the tracker's summary of the launch, not Microsoft's own price page, and the Flash latency is a vendor claim that nobody here has measured. Confirm both on Microsoft's page before you sign anything. Sume's figures are on its pricing page and in its TTS Router catalog at GET /v1/tts-router/models; see unknown-model errors for how the catalog answers.
A practical way to decide without a long evaluation: write down the one sentence that describes your worst-case user moment. If it is "the customer is on the phone and the agent has to answer", you need a live voice and the question is closed. If it is "the editor presses Generate and reviews the take", then price, voice-language checks and the pieces around the file matter more than the first-byte time, and a job API is the simpler thing to run. Most voice features that ship in a product turn out to be the second kind.
If you want to test the Sume side, time a job yourself: the latency benchmark script takes a few minutes and tells you whether an async file is fast enough for your flow.
Sources
Related posts
More in Comparisons
- MiniMax H3 open weights vs a hosted lip-sync API: what you take on
MiniMax released H3 open weights on 2026-08-03. Self-hosting is not the same as a hosted still-plus-audio lip-sync route. What each choice makes you own.
- Nova Canvas 4,194,304 pixel cap and 16 px rule vs Sume image_size
Nova Canvas output sides must divide by 16 and total under 4,194,304 pixels. GPT Image 2.5 on Sume has its own custom-size rules. A side-by-side check.
- Nova Canvas IMAGE_VARIATION similarityStrength vs Sume n and reference
Nova Canvas IMAGE_VARIATION takes 1 to 5 images and a similarityStrength of 0.2 to 1.0. Sume has no strength field; here is how to get variations.
- Nova Canvas inpainting mask rules vs a Sume reference edit
Nova Canvas inpainting wants a pure black and white mask, same size as the input. What to build, and how Sume's edit call differs.
Written by Sume