Voice API deadlines, October 2026 to February 2027
A calendar of voice and transcription API changes from vendor pages: Gemini TTS price rise, OpenAI transcription shutdown, and the xAI voice alias move.

Two dates matter for anyone with a voice integration this winter. Gemini's 3.8 Flash TTS prices double on 1 January 2027, and OpenAI's whisper-1 and three gpt-4o transcription models shut down on 26 February 2027. xAI already moved its grok-voice-latest alias on 5 August 2026. Check the calendar before you ship another release.
The calendar
| Date | Vendor | Change |
|---|---|---|
| 5 Aug 2026 (past) | xAI | grok-voice-latest moved to Think Fast 2.0; users who had not pinned were upgraded |
| 26 Aug 2026 | OpenAI | Deprecation notice: whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize |
| 31 Dec 2026 | Gemini 3.8 Flash TTS at $0.50 text input and $9.00 audio output per 1M tokens ends | |
| 1 Jan 2027 | Flash TTS becomes $1.00 and $18.00; Flash-Lite TTS $1.00 and $12.00 (from $0.50 and $6.00) | |
| 26 Feb 2027 | OpenAI | Shutdown of the four deprecated transcription models; move to gpt-live-transcribe or gpt-transcribe |
What to do on each
- Pin explicit model versions today. Aliases that say latest move without warning.
- Budget the January TTS price step now if you use Gemini TTS at volume.
- Run your transcription test set on the replacement model before February, not in the last week.
The same discipline on Sume
Sume's TTS router has sonic-latest next to explicit ids such as sonic-3.6; pin an explicit id if you do not want output to change under you. Music 1.0 is retiring gradually and resolves through the music router, so new work should call the router. The API reference lists the live model ids, and the catalog endpoint shows what is available now.
Limits
Vendors reschedule. These rows are what each page said on 1 October, and they will age. Prices are quoted from the vendor pages, not verified by a test call.
Sources
More in Developers
- Which Sume audio endpoint to call: TTS, STT, music, detach, timeline
A decision map for Sume's audio API: seven endpoints, what each takes in and returns, limits and list prices, and the order they chain in.
- Sume video tools: public URL or media import first? Per tool
Video captions takes a public HTTPS URL; trim, filter, inspect, frames, compose and detach need a workspace media.sume.com clip. A tool-by-tool input guide.
- Which voice does my avatar speak with? Check voice.status is ready
Sume TTS speaks in an avatar's voice when voice.status is ready. List avatars, check voice.status, then send avatar_id or avatar_handle on the TTS request.
- YouTube captions.insert: 100 MB, 400 quota units, and an SRT build
YouTube captions.insert costs 400 quota units and takes a 100 MB file. Sume returns words and segments, not SRT, so here is the 20-line conversion to upload.
Written by Sume