Microsoft's October 2026 speech launches: what to change in a pipeline

MAI-Voice-2.1, its Flash model and MAI-Transcribe-2-Streaming, claim by claim, with what each means for an ad voice pipeline and what Sume covers today.

5 min readSume
All posts

Microsoft's October 2026 speech launches change three decisions in an ad voice pipeline: which model to use when latency matters, how to treat voice consent, and whether to transcribe live. Sume covers recorded voiceover, transcription and captions as asynchronous jobs, and it does not offer a streaming voice or live transcription.

Claim by claim

Every vendor number below was read from Microsoft's own pages on 2026-10-05, and the Sume column comes from the Sume docs and API reference. Where the pages say nothing, the table says so.

Microsoft speech claims and what they mean for a Sume pipeline (vendor pages and Sume docs, read 2026-10-05)
LaunchVendor claimWhat to doOn Sume
MAI-Voice-2.1 StandardAbout 550 ms latency; $22 per 1M characters; 23 languages; emotion controlTest it for long, recorded content, where its page says audiobooks and content fitTTS 1.0 and TTS Router, $0.0475 per 1,000 characters, per job
MAI-Voice-2.1 FlashAbout 45 ms latency; $15 per 1M characters; call centres and assistantsTreat as a live-voice option; test from your own regionNo streaming TTS; jobs are async, sync wait at most 30 s
MAI-Voice-2 licensingOnly authorised, licensed voices in production; no unlicensed cloningWrite the same rule into your own voice policyUse voices you hold rights to
MAI-Transcribe-2-StreamingReal-time transcription; 60 languages; automatic detection; diarisation; timestampsUse for live captions and callsRecorded audio only: STT 1.0 at $0.01 a minute, auto-detect when language_code is omitted

Decision one: latency

Microsoft's model page puts Standard at roughly 550 ms and Flash at roughly 45 ms. A tracker reports a different Flash figure, 150 ms end to end, which is a reminder that a number depends on what is measured and from where. Use your own timing, taken from your own region with your own text.

If the voice has to answer a person while they wait, a streaming engine is the right tool, and Sume does not ship one: streaming TTS is a stated non-goal for the current TTS release. If the voice is for an ad that a team reviews before it airs, latency is irrelevant and quality and consent matter more.

Decision two: consent

Microsoft's announcement says only authorised, licensed voices can be synthesised in production and that unlicensed cloning is not possible. That is a policy position worth copying into your own rules, whichever engine you use. Before a first take, an ad team should know whose voice it is, what the licence allows, and where the paper lives.

Decision three: transcription

A streaming transcriber with 60 languages is aimed at live meetings and calls. For a folder of finished videos, the Sume path is different: detach the audio ($0.01), transcribe at $0.01 per minute, and burn captions at $0.20 per clip of up to 60 seconds. A 9-minute recording costs $0.10 to transcribe. There is no live path.

  • Live captions on a call or a stream: use a streaming service.
  • Subtitles on an ad you are about to publish: use the recorded path.
  • Language unknown: omit language_code and let STT detect it.

What does not change

Your script, your sentence count and your retake budget. A 90-second read is about 1,350 characters under a 15 characters a second assumption, and on Sume it is 7 cents. On Microsoft's published prices the same text is about 2 to 3 cents before rounding. At this length the engine is not the cost driver; the render and the captions are.

The unchanged work is listening. Run the same three test lines, a brand name, a number and an emotional line, through each candidate and decide with your ears.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume