Deepgram self-hosted drops Whisper in 261001: a hosted STT job instead

Deepgram's release 261001 removes Whisper support from self-hosted. Your options: move to Nova-3, or send recordings to a hosted STT job such as Sume's.

5 min readSume
All posts

Deepgram's self-hosted release 261001, published October 1, 2026, removes Whisper model support, and the changelog tells Whisper users to migrate to Nova-3. If you ran Whisper through Deepgram's containers, you have three choices: move to Nova-3 on the same stack, pin the older release, or send recordings to a hosted speech-to-text job. This post covers the third option using Sume STT, and is honest about where it does not fit.

Deepgram details are from its changelog (read 2026-10-03). Sume details come from the OpenAPI reference and the jobs docs.

What does release 261001 change for Whisper users?

The release publishes four images tagged release-261001: the API, the engine, the license proxy and billing, with -fips variants. The engine needs an NVIDIA driver version 580 or newer. The headline for this audience is that Whisper support is removed and Nova-3 is the stated migration target. The same release makes Flux TTS watermarking mandatory and improves Japanese numerals and German number reading.

The changelog does not say how long older releases stay available or supported, so do not assume you can stay on a previous tag indefinitely. Ask Deepgram support before you plan around a pin.

What are your options?

The right answer depends on why you chose Whisper in the first place: language coverage, cost per hour on your own GPUs, or a data rule that audio cannot leave your network. Only the last one rules out a hosted job.

Options after Whisper removal in Deepgram self-hosted 261001 (read 2026-10-03)
OptionFits whenCost of the move
Nova-3 on your self-hosted stackAudio must stay on your networkRe-test accuracy and formatting on your recordings
Stay on an older releaseYou need a short bridgeUnknown support window; ask the vendor
Hosted STT job such as Sume STTAudio can leave your networkUpload to a public HTTPS URL; per-minute billing

What does a Sume STT job look like for former Whisper traffic?

Sume STT is POST /v1/stt-1.0/transcribe with a public HTTPS audio_url. You can send a language_code hint or omit it for auto-detect, a duration_seconds between 1 and 600 so usage is reserved accurately, and segmentation of mode sentence if you want sentence groups built from word timings. Word timings are always returned as words[], in seconds from the start of the audio, so you do not need a flag to get them.

The limits matter for batch workloads. A job is capped at ten minutes of audio, a sync wait lasts at most 30 seconds, and diarization and similar knobs are fixed server-side. Price is $0.01 per audio minute at the public rate, so a one-hour recording split into six ten-minute parts is about sixty cents before any splitting work.

How do you handle long recordings?

Whisper-based pipelines often feed hour-long files. With a ten-minute ceiling per job you split first. Sume's audio detach can pull the audio track from a video, and timeline audio can split a file into up to 20 ranges per job, each with its own hosted URL. Then submit one STT job per range with async mode, poll the statuses, and stitch the transcripts by adding each range's start offset to its word timings.

Keep the idempotency key per range stable, so a retry after a network error returns the same job rather than billing a second time.

  • Split into ranges under ten minutes, with a little overlap at the seams.
  • Submit one STT job per range in async mode.
  • Offset each range's word times by its start time before merging.
  • Drop duplicate words in the overlap, then keep the job ids next to the merged transcript.

What should you test before switching?

Run a sample of your real recordings through both the Nova-3 route and a hosted job, and compare word error on the terms that matter to you: product names, numbers and speaker accents. Check that the output format your downstream code expects, such as sentence segments or timings, is present on the route you pick.

If the recordings are sensitive, read the hosted vendor's terms before you upload. Sume's docs describe workspace-scoped jobs and hosted artifacts, but they do not give a data-processing location, which can decide the question on its own.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume