TTS price per 1,000 characters after Oct 12: ranking incl. Sume

Ranking list TTS prices per 1,000 characters for the week of Oct 13: MAI, OpenAI, ElevenLabs, Deepgram and Sume. Where Sume sits, and what the list leaves out.

4 min readSume
All posts

After ElevenLabs' launch discount ends on October 12, the cheapest list prices I read per 1,000 characters are $0.015 for Microsoft's MAI-Voice-2.1-Flash and OpenAI's tts-1, and the most expensive is Eleven v4 at $0.08. Sume's TTS Router sits near the middle at about $0.0475, just above Deepgram Flux TTS at $0.045. A price ranking is a starting point: the engines sound different, and some are tuned for live calls.

Gemini TTS is left out because it is priced per million tokens, not per character, and I did not find a conversion on the pages I read.

Ranking, cheapest first

Prices are as read on the pages named in Sources. Sume's figure is Cartesia list times 1.25, rounded up to the cent per job.

List price per 1,000 characters (read 2026-10-08)
RankServicePriceNote
1 (tie)Microsoft MAI-Voice-2.1-Flash$0.015About 150 ms end to end for up to 45 s
1 (tie)OpenAI tts-1$0.015Per OpenAI pricing page
3Microsoft MAI-Voice-2.1$0.02223 languages, 26 locales
4OpenAI tts-1-hd$0.030Per OpenAI pricing page
5ElevenLabs v4 Turbo (list)$0.04Launch discount ends Oct 12
6Deepgram Flux TTS$0.045English only; for live agents
7Sume TTS Router (Sonic)About $0.0475Async jobs, up to 20,000 characters
8ElevenLabs v4 (list)$0.08Launch discount ends Oct 12

Use the ranking in four steps

  • Remove any engine that lacks a language, a voice or a delivery mode you need.
  • Listen to the same script on the three cheapest remaining.
  • Multiply the price by your monthly characters.
  • Re-check on the first of each month, since prices and discounts move.
import os, requests

r = requests.post(
    "https://api.sume.com/v1/tts-router/generate",
    headers={
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": "tts-demo-001",
    },
    json={
        "model": "sonic-3.6",
        "transcript": "Welcome back. Today we compare three prices.",
        "voice": {"id": os.environ["SUME_VOICE_ID"]},
        "timestamps": {"words": True},
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

What the ranking leaves out

It ignores quality, language coverage, voice cloning and delivery mode. It also ignores minimums: Sume rounds each job up to the cent, so many tiny jobs cost more than one long job. A 50-character job costs a cent where a pure per-character price would be a fraction of one. Batch your short lines if cost matters.

How to read ties and small gaps

The first two rows tie at $0.015 and yet serve different needs: Microsoft lists its Flash model for low-latency use with about 150 ms end to end for up to 45 seconds of audio, while OpenAI's tts-1 is a general model. A one-cent difference per 1,000 characters is 10 dollars per million, which is rarely the deciding factor once you hear the voices side by side.

The same logic applies to Sume at about $0.0475 against Deepgram at $0.045. A gap of 0.25 cents per 1,000 characters is $2.50 per million characters. Voice fit, language and delivery mode should decide, and the price should break the tie.

What Sume does not do

Sume does not offer the other engines in this table. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. Prices for other vendors may differ by plan, region and contract. I read these as published list prices.

Related

See Eleven v4 at $80 per million characters and Gemini TTS after the January rise.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume