grok-voice-transcribe-1.0 ended Oct 2: what a silent reroute means
xAI ended grok-voice-transcribe-1.0 on Oct 2, 2026 and routes it to 2.0 at the same price. How to catch silent model swaps, and what Sume STT 1.0 fixes for you.

xAI's release notes say grok-voice-transcribe-1.0 reached end of life on October 2, 2026, and that requests to that slug are now routed to grok-voice-transcribe-2.0 at the same price, with improved accuracy. Your code keeps working, but the model behind the name changed, so a transcript you cached in September may differ from one you produce today. The fix is to record which model answered, and to compare before you trust the swap.
xAI facts below are from its release notes, read on 2026-10-03. Sume facts come from the OpenAPI reference.
What did xAI change on October 2?
Per the release notes, the 1.0 slug no longer runs 1.0. Calls to it return results from 2.0, the price is unchanged, and xAI describes the new model as more accurate. The same page lists earlier speech-to-text changes: Speech to Text reached general availability in April 2026 with 25 languages in batch and streaming modes, and a July 2026 update added a vad_threshold parameter to tune the voice-activity gate that skips non-speech audio, plus Smart Turn for streaming transcription.
Nothing in the notes says the response shape changed. What changes is the text: different punctuation, different spellings of names, sometimes different word boundaries. Anything that keyed on exact text, such as a golden test, a cache or a subtitle diff, can break without an error.
| Date | Item |
|---|---|
| April 2026 | Speech to Text generally available, 25 languages, batch and streaming |
| July 2026 | vad_threshold parameter; Smart Turn for streaming transcription |
| October 2, 2026 | grok-voice-transcribe-1.0 ends; requests route to 2.0 at the same price |
How do you detect a silent model swap?
Three habits catch it. First, store the model id you requested alongside every transcript, and any model field the response returns, so a later diff has something to explain it. Second, keep a small regression set of recordings with known-good transcripts and re-run it on a schedule, not only when you deploy. Third, alert on cost and word-count drift per hour of audio, since a model change often moves both.
A reroute at the same price is the friendly case. When a slug is retired and routed to a pricier or differently shaped model, the same habits turn a billing surprise into a ticket you saw coming.
What does Sume STT 1.0 pin for you?
Sume's route is POST /v1/stt-1.0/transcribe, and callers use the public model id sume/stt-1.0; the reference says provider model ids stay internal. That gives you one name in your code and one job record per request. It does not promise the provider behind it never changes, and the docs do not publish a model-change log, so the regression set above still applies.
The request itself is small: a public HTTPS audio_url, an optional language_code hint, an optional duration_seconds from 1 to 600, optional sentence segmentation, and metadata that is stored on the job and not sent to the provider. Word timings come back as words[] in seconds from the start. Price is $0.01 per audio minute, with one minute reserved when duration_seconds is omitted and a ten-minute ceiling per job.
Which controls does each route give you?
If you need to tune the voice-activity gate or use Smart Turn on a live stream, xAI's API exposes those and Sume's does not: the reference says diarization and similar provider knobs are fixed server-side. If you want a stable Sume-owned job envelope for recorded audio, with status_url, result_url and a cancel path, Sume's route fits, and you can store the job id next to each transcript.
A short migration checklist for any speech-to-text slug that is being retired:
- Search your code and config for the old slug, including environment variables and workflow files.
- Re-run your regression recordings on the new model and diff words, not only totals.
- Record the model id and the date next to each transcript you keep.
- Check the retired model's end-of-life date against any reserved-capacity or contract terms.
Sources
Related posts
More in Comparisons
- Ideogram 4.5 Magic Fill and Extend vs Sume mask_url and aspect ratio
Ideogram's docs list Magic Fill and Extend as editing features. Sume has no Extend endpoint; here is what its mask and aspect-ratio fields cover instead.
- Instagram's Devanagari caption fonts: what Sume burns for Hindi
Instagram added Devanagari and Bengali-Assamese caption fonts. Sume documents Latin and Hangul faces only, so test Hindi cues on one clip first.
- Reels run to 20 minutes; Sume's audio and caption limits to plan for
Instagram says Reels can reach 20 minutes but over 3 minutes is not recommended. The Sume Timeline audio, avatar and caption limits that matter at that length.
- Kling 3 vs Seedance 2.5 vs Wan 3.0: 5-second 720p price on Sume
A 5-second 720p clip on Sume: Wan 3.0 $0.625, Kling 3 with audio $1.05 (audio off $0.70), Seedance 2.5 $2.889. List x 1.25 for every model.
Written by Sume