Microsoft's October 2026 speech launches: what to change in a pipeline
MAI-Voice-2.1, its Flash model and MAI-Transcribe-2-Streaming, claim by claim, with what each means for an ad voice pipeline and what Sume covers today.

Microsoft's October 2026 speech launches change three decisions in an ad voice pipeline: which model to use when latency matters, how to treat voice consent, and whether to transcribe live. Sume covers recorded voiceover, transcription and captions as asynchronous jobs, and it does not offer a streaming voice or live transcription.
Claim by claim
Every vendor number below was read from Microsoft's own pages on 2026-10-05, and the Sume column comes from the Sume docs and API reference. Where the pages say nothing, the table says so.
| Launch | Vendor claim | What to do | On Sume |
|---|---|---|---|
| MAI-Voice-2.1 Standard | About 550 ms latency; $22 per 1M characters; 23 languages; emotion control | Test it for long, recorded content, where its page says audiobooks and content fit | TTS 1.0 and TTS Router, $0.0475 per 1,000 characters, per job |
| MAI-Voice-2.1 Flash | About 45 ms latency; $15 per 1M characters; call centres and assistants | Treat as a live-voice option; test from your own region | No streaming TTS; jobs are async, sync wait at most 30 s |
| MAI-Voice-2 licensing | Only authorised, licensed voices in production; no unlicensed cloning | Write the same rule into your own voice policy | Use voices you hold rights to |
| MAI-Transcribe-2-Streaming | Real-time transcription; 60 languages; automatic detection; diarisation; timestamps | Use for live captions and calls | Recorded audio only: STT 1.0 at $0.01 a minute, auto-detect when language_code is omitted |
Decision one: latency
Microsoft's model page puts Standard at roughly 550 ms and Flash at roughly 45 ms. A tracker reports a different Flash figure, 150 ms end to end, which is a reminder that a number depends on what is measured and from where. Use your own timing, taken from your own region with your own text.
If the voice has to answer a person while they wait, a streaming engine is the right tool, and Sume does not ship one: streaming TTS is a stated non-goal for the current TTS release. If the voice is for an ad that a team reviews before it airs, latency is irrelevant and quality and consent matter more.
Decision two: consent
Microsoft's announcement says only authorised, licensed voices can be synthesised in production and that unlicensed cloning is not possible. That is a policy position worth copying into your own rules, whichever engine you use. Before a first take, an ad team should know whose voice it is, what the licence allows, and where the paper lives.
Decision three: transcription
A streaming transcriber with 60 languages is aimed at live meetings and calls. For a folder of finished videos, the Sume path is different: detach the audio ($0.01), transcribe at $0.01 per minute, and burn captions at $0.20 per clip of up to 60 seconds. A 9-minute recording costs $0.10 to transcribe. There is no live path.
- Live captions on a call or a stream: use a streaming service.
- Subtitles on an ad you are about to publish: use the recorded path.
- Language unknown: omit
language_codeand let STT detect it.
What does not change
Your script, your sentence count and your retake budget. A 90-second read is about 1,350 characters under a 15 characters a second assumption, and on Sume it is 7 cents. On Microsoft's published prices the same text is about 2 to 3 cents before rounding. At this length the engine is not the cost driver; the render and the captions are.
The unchanged work is listening. Run the same three test lines, a brand name, a number and an emotional line, through each candidate and decide with your ears.
Sources
Related posts
More in Comparisons
- Midjourney edits now touch only selected pixels; the API route on Sume
Midjourney's Sep 24 update says edits modify only selected pixels. Midjourney isn't on Sume; a masked edit via mask_url on ChatGPT Image 2.5 is the API route.
- Midjourney V8.1 draft mode makes 24 images; Sume caps n at 10 per call
Midjourney V8.1 draft mode returns 24 lower-res images per job. Sume's n is capped 1-10 per request and lower per model, so 24 images takes 3 to 6 calls.
- Nano Banana 2 Lite alternative on Sume: image rows under 4 cents
Sume lists no Nano Banana 2 Lite row. Seven rows quote under $0.04, from GPT Image 2.5 low at $0.0074 to Flux 2 Pro at $0.0375. Table with edit support.
- Nano Banana 2 vs Ideogram 4.5 for photo edits: ratios, 4K, refs
Nano Banana 2 offers 512 to 4K, 21:9 and 10 references on Sume. Ideogram 4.5 offers 1K and 2K, 15 ratios, 5 references and keeps the source shape on edits.
Written by Sume