Deepgram self-hosted drops Whisper in 261001: a hosted STT job instead
Deepgram's release 261001 removes Whisper support from self-hosted. Your options: move to Nova-3, or send recordings to a hosted STT job such as Sume's.

Deepgram's self-hosted release 261001, published October 1, 2026, removes Whisper model support, and the changelog tells Whisper users to migrate to Nova-3. If you ran Whisper through Deepgram's containers, you have three choices: move to Nova-3 on the same stack, pin the older release, or send recordings to a hosted speech-to-text job. This post covers the third option using Sume STT, and is honest about where it does not fit.
Deepgram details are from its changelog (read 2026-10-03). Sume details come from the OpenAPI reference and the jobs docs.
What does release 261001 change for Whisper users?
The release publishes four images tagged release-261001: the API, the engine, the license proxy and billing, with -fips variants. The engine needs an NVIDIA driver version 580 or newer. The headline for this audience is that Whisper support is removed and Nova-3 is the stated migration target. The same release makes Flux TTS watermarking mandatory and improves Japanese numerals and German number reading.
The changelog does not say how long older releases stay available or supported, so do not assume you can stay on a previous tag indefinitely. Ask Deepgram support before you plan around a pin.
What are your options?
The right answer depends on why you chose Whisper in the first place: language coverage, cost per hour on your own GPUs, or a data rule that audio cannot leave your network. Only the last one rules out a hosted job.
| Option | Fits when | Cost of the move |
|---|---|---|
| Nova-3 on your self-hosted stack | Audio must stay on your network | Re-test accuracy and formatting on your recordings |
| Stay on an older release | You need a short bridge | Unknown support window; ask the vendor |
| Hosted STT job such as Sume STT | Audio can leave your network | Upload to a public HTTPS URL; per-minute billing |
What does a Sume STT job look like for former Whisper traffic?
Sume STT is POST /v1/stt-1.0/transcribe with a public HTTPS audio_url. You can send a language_code hint or omit it for auto-detect, a duration_seconds between 1 and 600 so usage is reserved accurately, and segmentation of mode sentence if you want sentence groups built from word timings. Word timings are always returned as words[], in seconds from the start of the audio, so you do not need a flag to get them.
The limits matter for batch workloads. A job is capped at ten minutes of audio, a sync wait lasts at most 30 seconds, and diarization and similar knobs are fixed server-side. Price is $0.01 per audio minute at the public rate, so a one-hour recording split into six ten-minute parts is about sixty cents before any splitting work.
How do you handle long recordings?
Whisper-based pipelines often feed hour-long files. With a ten-minute ceiling per job you split first. Sume's audio detach can pull the audio track from a video, and timeline audio can split a file into up to 20 ranges per job, each with its own hosted URL. Then submit one STT job per range with async mode, poll the statuses, and stitch the transcripts by adding each range's start offset to its word timings.
Keep the idempotency key per range stable, so a retry after a network error returns the same job rather than billing a second time.
- Split into ranges under ten minutes, with a little overlap at the seams.
- Submit one STT job per range in async mode.
- Offset each range's word times by its start time before merging.
- Drop duplicate words in the overlap, then keep the job ids next to the merged transcript.
What should you test before switching?
Run a sample of your real recordings through both the Nova-3 route and a hosted job, and compare word error on the terms that matter to you: product names, numbers and speaker accents. Check that the output format your downstream code expects, such as sentence segments or timings, is present on the route you pick.
If the recordings are sensitive, read the hosted vendor's terms before you upload. Sume's docs describe workspace-scoped jobs and hosted artifacts, but they do not give a data-processing location, which can decide the question on its own.
Sources
Related posts
More in Comparisons
- DeepL Voice keeps speakers' voices in 14 languages: Sume dubs
DeepL's September 15 release keeps each speaker's voice across 14 languages in live talk. A Sume dub picks a TTS voice id per line and does not clone.
- Does Sume have Veo 3.1? No. Which video models it lists instead
Sume's video catalog has no Veo model. See the models it does list, with duration and resolution ranges, and how Google's own docs now steer to Omni Flash.
- Eleven v4: 90+ languages or 99? ElevenLabs' own pages differ
ElevenLabs' launch post says more than 90 languages for Eleven v4; its docs page lists 99. How to plan around the gap, and what Sume lists instead.
- Eleven v4 drops the native accent; Sume keeps a language tag per voice
ElevenLabs' docs say v4 gives fluent target-language speech, not a preserved accent. Sume tags each voice with a language and checks for mismatch.
Written by Sume