AI voice news, week of Oct 7, 2026: what changed for voiceover teams
Microsoft MAI-Voice-2.1, Decagon Chord and Voice 3, Gemini 3.8 TTS price step and Voxtral TTS: a dated list of what to act on this week.

Three things changed for teams that make voiceovers in the last two weeks: Microsoft added MAI-Voice-2.1 and a Flash variant, Decagon launched Voice 3 with its Chord speech model for live calls, and Google shipped Gemini 3.8 TTS, whose paid prices double on January 1, 2027. Mistral's Voxtral TTS (March 2026) is the open-weights reference point. Here is the dated list.
What happened and when
Each row is from the vendor's own page.
| Date | Vendor | What | Why it matters |
|---|---|---|---|
| Sep 23, 2026 | Gemini 3.8 Flash TTS and Flash-Lite TTS | Over 100 languages; voice design; paid output prices double on Jan 1, 2027 | |
| Oct 1, 2026 | Microsoft | MAI-Voice-2.1, Flash, MAI-Transcribe-2-Streaming | 23 languages (public preview); transcription $0.54 per hour introductory through year end |
| Oct 1, 2026 | Decagon | Voice 3 with Chord | Duplex live agent; 70+ languages |
| Mar 2026 | Mistral | Voxtral TTS | 9 languages; $0.016 per 1,000 characters; weights under CC BY-NC 4.0 |
Actions by role
Give each person one thing to do.
- Producers: listen to one script on MAI-Voice-2.1, Gemini Flash TTS and your current voice, blind.
- Finance: add Gemini's January 1, 2027 price step and Microsoft's end-of-year intro rate to the 2027 budget.
- Developers: note that Microsoft marks the MAI voices as public preview.
- Support leads: if you run live calls, read the Voice 3 page.
Where Sume stands
Sume's TTS 1.0 is $0.0475 per 1,000 characters and speech-to-text is $0.01 per audio minute on the public rate card, with word timings and sentence slices in the result. Sume does not list the vendors above as models on its TTS surface, so use their files by importing them, or use Sume's own voices. Confirm rates in GET /v1/catalog.
What not to do
Do not switch a live client campaign to a preview voice in the week it launches. Keep the approved files, test on a new draft, and decide after a listening round.
Sources
- Microsoft AI: Our first streaming transcription model (read 2026-10-07)
- Microsoft Learn: MAI-Voice-2.1 and MAI-Voice-2.1-Flash (read 2026-10-07)
- Decagon: Introducing Voice 3 and Chord (read 2026-10-07)
- Google: Gemini 3.8 text-to-speech says hello (read 2026-10-07)
- Google AI: Gemini API pricing (read 2026-10-07)
- Mistral: Voxtral TTS (read 2026-10-07)
- Sume API pricing
Related posts
More in Models
- Anime-style ad with Seedance 2.5: use a character sheet as references
Send your mascot's character sheet as input_references to seedance-2.5 on Sume and test a 6-second anime-style ad at 480p before you pay for 720p.
- ByteDance models on Sume: Seedance family and Seedream prices
Seedance 2.5, 2.0, Fast and Mini plus Seedream 5 Lite, 4.5 Edit and v4, with Sume prices, limits and the pick for each job. Read 2026-10-07.
- Can Gemini Omni Flash speak my script? It takes no audio input on Sume
On Sume, Gemini Omni Flash 1.1 takes image and video references but no audio reference. It makes its own sound. For your script in a face, use avatar video.
- Same character in three AI video shots: Omni reference image
Keep one character across separate Gemini Omni clips by sending the same reference image with every request on Sume. Request shape, prompt tags, 3-shot cost.
Written by Sume