MAI-Voice-2.1 or Sume TTS? Three questions that decide it

Live speech, extra outputs or lowest price per character? A short guide to MAI-Voice-2.1 and Flash against the Sume TTS Router, with 10M-character math.

5 min readSume
All posts

Choose Microsoft's MAI-Voice-2.1 Flash if speech has to start inside a live turn, and MAI-Voice-2.1 if per-character price decides and you do not need anything else from the API. Choose Sume's TTS Router if you want a finished audio file that comes with word timings, sentence segments, a voice-language guard and signed webhooks on the same job, and if the same account also covers captions, lip sync and video. Three questions get you there. The facts below come from the October 2026 tracker (read 2026-10-06) and Sume's pricing page.

1. Does speech start inside a live turn?

The tracker lists MAI-Voice-2.1-Flash at a vendor-claimed 150 ms end to end. Sume TTS is an async job that returns a hosted file, and streaming TTS is a documented non-goal of the router. If a person is waiting for the first word, that settles it for the Microsoft product. If a person presses a button and waits a few seconds for a clip, it does not. Sync mode waits at most 30 seconds on one request and falls back to polling, so it fits a button, not a call.

2. How much of the bill is characters?

MAI-Voice-2.1 is listed at $22 per 1M characters and Flash at $15. Sume's router is $47.50 per 1M on every model, the provider list price with a 1.25 margin. At small volume the gap is cents. At 10M characters a month it is $220 against $475, or $150 for Flash. If the characters are the whole bill, that is the number to weigh. If you also pay for transcription, caption renders, a second vendor's webhooks and the glue between them, count those too.

3. What else do you need next to the audio?

A voiceover is rarely the final product. Sume's TTS job can return words[] with start and end times, and gapless sentence segments[] with a 70 ms default boundary lead. The completed job records the model, voice, language and output settings. A wrong-language voice stops with a 409 before it spends. The same account runs speech-to-text at $0.01 per audio minute, captions at $0.20 per job and Timeline audio at $0.01 per job. Whether that matters is a question about your pipeline, not about the voice.

Decision table (read 2026-10-06)
If you needMAI-Voice-2.1MAI-Voice-2.1-FlashSume TTS Router
Speech inside a live turnNot statedVendor claims 150 ms end to endNo streaming
Lowest characters bill, 10M a month$220$150$475
Word timings and sentence segments with the audioNot stated in the trackerNot stated in the trackerYes, one request
Languages23 languages, 26 localesNot stated separatelySet language, guard checks the voice
Captions, lip sync and video on one keyNot in the trackerNot in the trackerYes

Use both, and check the numbers

Many teams end up with both: a streaming model for the live path and Sume for everything that ends as a file. Keep the boundary clean. Decide per surface, not per company, and put the choice behind one function so you can swap a voice without touching the caller.

One caution about the table: the Microsoft figures are the tracker's summary of the launch, not Microsoft's own price page, and the Flash latency is a vendor claim that nobody here has measured. Confirm both on Microsoft's page before you sign anything. Sume's figures are on its pricing page and in its TTS Router catalog at GET /v1/tts-router/models; see unknown-model errors for how the catalog answers.

A practical way to decide without a long evaluation: write down the one sentence that describes your worst-case user moment. If it is "the customer is on the phone and the agent has to answer", you need a live voice and the question is closed. If it is "the editor presses Generate and reviews the take", then price, voice-language checks and the pieces around the file matter more than the first-byte time, and a job API is the simpler thing to run. Most voice features that ship in a product turn out to be the second kind.

If you want to test the Sume side, time a job yourself: the latency benchmark script takes a few minutes and tells you whether an async file is fast enough for your flow.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume