Voice API cost gap up to 213x: what drives it, how to price yours

Krisp's newsletter reports a voice API cost gap of up to 213x across OpenAI, Gemini and Qwen. Why list TTS prices differ far less, and how to normalise yours.

4 min readSume
All posts

The 213x figure is a reported gap across voice APIs from OpenAI, Gemini and Qwen in Krisp's Voice AI newsletter, which gives no method I could verify. Across plain text-to-speech list prices I fetched this week, the spread is far smaller: from $0.015 to $0.08 per 1,000 characters is about 5.3x. A gap of 213x only appears when you compare different products, such as a realtime conversation model against a batch transcription tier.

So the useful move is to ignore the headline and put your own workload into one unit: cost per finished minute of audio.

List TTS prices, one unit

All rows are per 1,000 characters, from vendor pages or news reports as named in Sources. Sume's row is Cartesia list times 1.25, rounded up to the cent for each job.

TTS price per 1,000 characters (read 2026-10-08)
ServicePrice per 1,000 characters
Microsoft MAI-Voice-2.1-Flash$0.015
OpenAI tts-1$0.015
Microsoft MAI-Voice-2.1$0.022
OpenAI tts-1-hd$0.030
ElevenLabs v4 Turbo (list)$0.04
Deepgram Flux TTS$0.045
Sume TTS Router (Sonic)About $0.0475
ElevenLabs v4 (list)$0.08

Normalise your own price in four steps

  • Render three typical scripts and note the length of each in characters and the finished audio in minutes.
  • Divide characters by minutes to get your characters per finished minute.
  • Multiply by each service's price per character, or its token or minute equivalent.
  • Re-run the sheet when a price changes, and date every row.
chars_per_min = 900  # placeholder: replace with your measured value
price_per_1k = {
    "MAI-Voice-2.1-Flash": 0.015,
    "OpenAI tts-1": 0.015,
    "Deepgram Flux TTS": 0.045,
    "Sume TTS Router": 0.0475,
    "Eleven v4 (list)": 0.08,
}
for name, p in sorted(price_per_1k.items(), key=lambda kv: kv[1]):
    print(f"{name}: ${chars_per_min / 1000 * p:.4f} per finished minute")

What drives the big gaps

Three things stretch the range. Units differ: characters, audio tokens and minutes do not convert without a measured ratio. Products differ: a realtime model that listens and answers is not a narration engine. Tiers differ: batch, flex and priority modes can halve or raise a rate. Promotions add noise, such as ElevenLabs' 72% launch discount that ends on October 12.

A worked example with a placeholder

Say your narration runs 900 characters per finished minute. This number is a placeholder; measure yours. At that ratio, one finished minute costs about 1.4 cents on a $0.015 per 1,000 character price, about 4.3 cents on Sume's roughly $0.0475, and 7.2 cents on a $0.08 price. The spread between the cheapest and the dearest is about 5.3x, the same as the ratio of the prices, since the characters cancel out.

Seen this way, a 213x headline is a prompt to ask what was compared. If it contrasts a long realtime conversation with a short narration, the gap is real but irrelevant to your decision. The only gap that matters is between the engines that could do your job.

What Sume does not do

Sume does not offer OpenAI, Gemini or Qwen voices, and does not sell realtime conversation. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. I could not read the Krisp newsletter's underlying data, so the 213x figure is reported, not checked.

Related

See the price ranking after the Oct 12 promo and Qwen-Audio 3.1 price cuts.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume