Voice API cost gap up to 213x: what drives it, how to price yours
Krisp's newsletter reports a voice API cost gap of up to 213x across OpenAI, Gemini and Qwen. Why list TTS prices differ far less, and how to normalise yours.

The 213x figure is a reported gap across voice APIs from OpenAI, Gemini and Qwen in Krisp's Voice AI newsletter, which gives no method I could verify. Across plain text-to-speech list prices I fetched this week, the spread is far smaller: from $0.015 to $0.08 per 1,000 characters is about 5.3x. A gap of 213x only appears when you compare different products, such as a realtime conversation model against a batch transcription tier.
So the useful move is to ignore the headline and put your own workload into one unit: cost per finished minute of audio.
List TTS prices, one unit
All rows are per 1,000 characters, from vendor pages or news reports as named in Sources. Sume's row is Cartesia list times 1.25, rounded up to the cent for each job.
| Service | Price per 1,000 characters |
|---|---|
| Microsoft MAI-Voice-2.1-Flash | $0.015 |
| OpenAI tts-1 | $0.015 |
| Microsoft MAI-Voice-2.1 | $0.022 |
| OpenAI tts-1-hd | $0.030 |
| ElevenLabs v4 Turbo (list) | $0.04 |
| Deepgram Flux TTS | $0.045 |
| Sume TTS Router (Sonic) | About $0.0475 |
| ElevenLabs v4 (list) | $0.08 |
Normalise your own price in four steps
- Render three typical scripts and note the length of each in characters and the finished audio in minutes.
- Divide characters by minutes to get your characters per finished minute.
- Multiply by each service's price per character, or its token or minute equivalent.
- Re-run the sheet when a price changes, and date every row.
chars_per_min = 900 # placeholder: replace with your measured value
price_per_1k = {
"MAI-Voice-2.1-Flash": 0.015,
"OpenAI tts-1": 0.015,
"Deepgram Flux TTS": 0.045,
"Sume TTS Router": 0.0475,
"Eleven v4 (list)": 0.08,
}
for name, p in sorted(price_per_1k.items(), key=lambda kv: kv[1]):
print(f"{name}: ${chars_per_min / 1000 * p:.4f} per finished minute")What drives the big gaps
Three things stretch the range. Units differ: characters, audio tokens and minutes do not convert without a measured ratio. Products differ: a realtime model that listens and answers is not a narration engine. Tiers differ: batch, flex and priority modes can halve or raise a rate. Promotions add noise, such as ElevenLabs' 72% launch discount that ends on October 12.
A worked example with a placeholder
Say your narration runs 900 characters per finished minute. This number is a placeholder; measure yours. At that ratio, one finished minute costs about 1.4 cents on a $0.015 per 1,000 character price, about 4.3 cents on Sume's roughly $0.0475, and 7.2 cents on a $0.08 price. The spread between the cheapest and the dearest is about 5.3x, the same as the ratio of the prices, since the characters cancel out.
Seen this way, a 213x headline is a prompt to ask what was compared. If it contrasts a long realtime conversation with a short narration, the gap is real but irrelevant to your decision. The only gap that matters is between the engines that could do your job.
What Sume does not do
Sume does not offer OpenAI, Gemini or Qwen voices, and does not sell realtime conversation. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. I could not read the Krisp newsletter's underlying data, so the 213x figure is reported, not checked.
Related
See the price ranking after the Oct 12 promo and Qwen-Audio 3.1 price cuts.
Sources
- Krisp Voice AI newsletter (read 2026-10-08)
- OpenAI API pricing (read 2026-10-08)
- SiliconANGLE on MAI-Transcribe-2-Streaming (read 2026-10-08)
- Unite.AI on MAI-Transcribe-2 and MAI-Voice (read 2026-10-08)
- AudioXpress on Deepgram Flux TTS (read 2026-10-08)
- ElevenLabs API pricing (read 2026-10-08)
- Sume TTS Router catalog (read 2026-10-08)
- Sume API reference (read 2026-10-08)
Related posts
More in Pricing
- Voiceover cost for a 60-second video: about $0.043 on Sume TTS
A 60-second voiceover is roughly 900 characters, which bills about $0.043 on Sume TTS. The cost at 600 to 1,200 characters, and how to measure your own script.
- Wallet reserve for a full queue of Seedance 2.5 clips, by plan
Sume reserves the estimate at submit. A full queue of 5 s Seedance 2.5 720p clips ($2.889 each) reserves $17.33 on Free and $346.68 on Scale.
- Wan 3.0 prices for 10, 20 and 30 second clips: $0.63 to $7.50
Wan 3.0 on Sume costs $0.63 for 10 s at 480p and $7.50 for 30 s at 1080p. Nine prices, the per-second list rates and the rounding rule.
- Wan 3.0 on Sume: a 2-second minimum clip and a 30-second max, priced
Wan 3.0 accepts 2 to 30 seconds. On Sume a 2 s clip is $0.125 at 480p and $0.50 at 1080p; 30 s at 1080p is $7.50.
Written by Sume