New speech models, October 2026: which ones a video maker can use
MAI-Voice-2.1, MAI-Transcribe-2-Streaming, Pocket TTS and Mercury Voice, sorted by what a short-video maker can run today, with the Sume step that matches each.

Four speech releases landed around October 1, 2026, and only some matter for short video. Microsoft's MAI-Voice-2.1 and MAI-Transcribe-2-Streaming are hosted speech models, Kyutai's Pocket TTS is a small open model you run yourself, and Inception's Mercury Voice is a language model for voice agents, not a text-to-speech model. This is what each one is and the Sume step you would use for the same job.
What the vendors say
| Release | What it is | Stated numbers |
|---|---|---|
| MAI-Voice-2.1 | Hosted text to speech | 23 languages, 26 locales, $22 per million characters |
| MAI-Voice-2.1-Flash | Faster tier of the same | 150 ms end to end for 45 s of audio, $15 per million characters |
| MAI-Transcribe-2-Streaming | Hosted streaming speech to text | 60 languages, first partials just over 100 ms, $0.54 per hour through year end |
| Pocket TTS | Open-source TTS, 100M parameters | 7 languages, MIT license on GitHub, about 6x real time on a MacBook Air M4 |
| Mercury Voice | Diffusion LLM for voice agents | Median time to first token under 320 ms, enterprise GA October 1 |
What a short-video maker does with each
- MAI-Voice and Flash: voiceover. The Sume match is TTS at $0.0475 per 1,000 characters, and the render step ($0.10 per output minute) turns it into a 1080x1920 MP4.
- MAI-Transcribe-2-Streaming: live captions. A finished clip does not need streaming. Sume STT at $0.01 per audio minute gives word timings for captions on a file (API reference).
- Pocket TTS: a local voice for offline drafts. It is a model you host, with PyTorch 2.5 or newer, and its GitHub page notes it does not support silences or pauses in the text (GitHub).
- Mercury Voice: not a voiceover tool. It generates replies for a conversational agent (Inception Labs). Sume has no equivalent and this post does not claim one.
Rule of thumb
Pick by the shape of the job. Text in, file out: a batch TTS job. Audio in, text out, file already recorded: a batch STT job. Live talk with a person: a streaming model plus an agent LLM, which none of the Sume audio routes do. Only the first two are the daily work of a short-video team.
One caution on the numbers. Microsoft's $0.54 per hour is an introductory rate through the end of the year, and Inception's lower Mercury prices are a launch discount with no stated end date. Both can change, so recheck before you commit a budget.
Sources
Related posts
More in Models
- One voice in 23 languages: Microsoft's claim vs Sume's language guard
Microsoft says MAI-Voice-2.1 uses one voice across 23 languages. Sume sets language per job and warns on a voice mismatch. What that means when you buy.
- Pixal3D multi-view: prepare the input images
Pixal3D added multi-view inference in September 2026 under an MIT license. Sume has no 3D, but reference-image edit can prepare input views.
- PixVerse V6 native audio and camera work: what to check in an API
PixVerse's blog lists V6 with camera work and native audio, plus R2 and a $439M Series C total. How to test those claims against any video API's catalog.
- PixVerse V6 adds native audio: which Sume video ids make sound
PixVerse V6 ships camera work and native audio. The Sume video docs name no PixVerse id, so here is how to find models that generate audio and read the flag.
Written by Sume