Deepgram STT tiers per hour vs Sume: Nova-3, Whisper, Flux
Deepgram lists five STT rates from $0.0043 to $0.0065 a minute, or $0.258 to $0.39 an hour. Sume STT 1.0 is $0.01 a minute. The table prices 100 hours of each.

Deepgram's cheapest speech-to-text rate, Nova-3 pre-recorded mono at $0.0043 per minute, is $0.258 per hour, against $0.60 per hour for Sume STT 1.0. Its most expensive listed rate here, Flux streaming for English at $0.0065 per minute, is $0.39 per hour. Every Deepgram rate on the page read on 2026-10-05 is therefore lower than Sume's $0.01 per minute, by between 35 and 57 percent.
Five Deepgram rates and Sume
Deepgram's pricing page lists streaming and pre-recorded rates by model. The table picks five of them: three pre-recorded (Nova-3 mono, Whisper Large and Nova-3 multilingual) and two streaming (Nova-3 mono and Flux for English). It multiplies each per-minute rate by 60 for an hourly figure and by 6,000 for 100 hours. The last column is how much more Sume STT 1.0 would cost for the same 100 hours.
| Model | Rate | Per hour | 100 hours | Sume premium on 100 hours |
|---|---|---|---|---|
| Nova-3 pre-recorded, mono | $0.0043/min | $0.258 | $25.80 | $34.20 |
| Whisper Large pre-recorded | $0.0048/min | $0.288 | $28.80 | $31.20 |
| Nova-3 pre-recorded, multilingual | $0.0052/min | $0.312 | $31.20 | $28.80 |
| Nova-3 streaming, mono | $0.0048/min | $0.288 | $28.80 | $31.20 |
| Flux streaming, English | $0.0065/min | $0.390 | $39.00 | $21.00 |
| Sume STT 1.0 | $0.01/min | $0.600 | $60.00 | $0.00 |
Where the gap comes from
Sume STT 1.0 is the provider list of $0.008 per minute times 1.25, so $0.01 per minute, shown on the API pricing page. Deepgram's cheapest pre-recorded rate is below the provider list Sume starts from, so even before the margin the two are not the same model at the same price. The Sume figure is for a managed request type with one key and one balance, which is worth something if transcription is a step in a longer pipeline and worth nothing if it is the only thing you buy. Be clear with yourself about which of those you are.
What the table leaves out
Deepgram's page lists add-ons and different rates for other products, and the table uses mono audio. Streaming and pre-recorded work differently, and the mono assumption matters if your files are multichannel, because the page may price channels separately: the streaming rates suit live calls and the pre-recorded ones suit files. Sume STT 1.0 is a request-based, file-style service with a 10-minute limit per request, so the closest comparison is Deepgram's pre-recorded rows, which cost $0.258 to $0.312 an hour.
- Pre-recorded Nova-3 mono: 57 percent below Sume.
- Pre-recorded multilingual: 48 percent below Sume.
- Flux streaming: 35 percent below Sume.
Monthly volumes
Here are three monthly volumes for the nearest like-for-like pair: Deepgram Nova-3 pre-recorded mono and Sume STT 1.0. Each cell is hours times the hourly rate, with no free credits or plan fees.
The spread grows linearly, so the question is where it becomes larger than the cost of running a second vendor. For most teams that is somewhere in the hundreds of hours a month.
| Audio per month | Deepgram Nova-3 pre-recorded | Sume STT 1.0 | Difference |
|---|---|---|---|
| 50 hours | $12.90 | $30.00 | $17.10 |
| 500 hours | $129.00 | $300.00 | $171.00 |
| 5,000 hours | $1,290.00 | $3,000.00 | $1,710.00 |
When Sume is still the sensible choice
At 10 hours of audio a month the difference is $2.10 against the Flux streaming rate and $3.42 against Nova-3 pre-recorded, which no one picks a vendor for. At 1,000 hours it is $342 against the cheapest Deepgram rate. If your pipeline already generates video, speech and captions on Sume, the same balance covers the transcript. If transcription is the whole product, price it on Deepgram's page and run a 10-hour sample through each service first, using audio like your real traffic: the same languages, the same noise level and the same number of speakers. An accuracy gap of a few percentage points costs more in correction time than the per-hour gap saves.
Rates change, so read Deepgram's page the day you decide.
Sources
Related posts
More in Comparisons
- DeepSeek Flash or V4 Pro for a Sume job-polling loop: a price split
On DeepSeek's page Flash costs under a third of V4 Pro per token. Use it for jobs_wait polling; keep payload writing on one model. Read 2026-10-05.
- Draft vs final render cost: a 6-second clip on five Sume video models
The final render costs 2.5x to 10x the draft on Sume, depending on the model. A 6-second clip priced at the cheapest and the top tier, per model.
- Eleven v4 tags like [light rain] vs a separate sound bed on Sume
Eleven v4 puts effects such as [light rain] inside the speech. Sume keeps voice and bed separate, with gain_db, loop and duck_db. The trade-off for editing.
- Directing delivery: Eleven v4 audio tags vs Gemini TTS style
Eleven v4 puts direction inline as audio tags like [laughs]; Gemini TTS adds a separate style field. Sume sends transcript, voice and language only.
Written by Sume