New speech models, October 2026: which ones a video maker can use

MAI-Voice-2.1, MAI-Transcribe-2-Streaming, Pocket TTS and Mercury Voice, sorted by what a short-video maker can run today, with the Sume step that matches each.

4 min readSume
All posts

Four speech releases landed around October 1, 2026, and only some matter for short video. Microsoft's MAI-Voice-2.1 and MAI-Transcribe-2-Streaming are hosted speech models, Kyutai's Pocket TTS is a small open model you run yourself, and Inception's Mercury Voice is a language model for voice agents, not a text-to-speech model. This is what each one is and the Sume step you would use for the same job.

What the vendors say

Vendor-stated facts, each from the vendor's own page (read 2026-10-04)
ReleaseWhat it isStated numbers
MAI-Voice-2.1Hosted text to speech23 languages, 26 locales, $22 per million characters
MAI-Voice-2.1-FlashFaster tier of the same150 ms end to end for 45 s of audio, $15 per million characters
MAI-Transcribe-2-StreamingHosted streaming speech to text60 languages, first partials just over 100 ms, $0.54 per hour through year end
Pocket TTSOpen-source TTS, 100M parameters7 languages, MIT license on GitHub, about 6x real time on a MacBook Air M4
Mercury VoiceDiffusion LLM for voice agentsMedian time to first token under 320 ms, enterprise GA October 1

What a short-video maker does with each

  • MAI-Voice and Flash: voiceover. The Sume match is TTS at $0.0475 per 1,000 characters, and the render step ($0.10 per output minute) turns it into a 1080x1920 MP4.
  • MAI-Transcribe-2-Streaming: live captions. A finished clip does not need streaming. Sume STT at $0.01 per audio minute gives word timings for captions on a file (API reference).
  • Pocket TTS: a local voice for offline drafts. It is a model you host, with PyTorch 2.5 or newer, and its GitHub page notes it does not support silences or pauses in the text (GitHub).
  • Mercury Voice: not a voiceover tool. It generates replies for a conversational agent (Inception Labs). Sume has no equivalent and this post does not claim one.

Rule of thumb

Pick by the shape of the job. Text in, file out: a batch TTS job. Audio in, text out, file already recorded: a batch STT job. Live talk with a person: a streaming model plus an agent LLM, which none of the Sume audio routes do. Only the first two are the daily work of a short-video team.

One caution on the numbers. Microsoft's $0.54 per hour is an introductory rate through the end of the year, and Inception's lower Mercury prices are a launch discount with no stated end date. Both can change, so recheck before you commit a budget.

Sources

Related posts

More in Models

All Models posts

Written by Sume