OpenAI realtime mini ($10/$20 per M audio tokens) vs a narration job
OpenAI lists gpt-realtime mini at $10 in and $20 out per million audio tokens, and tts-1 at $15 per million characters. What each suits, and where Sume fits.

OpenAI's pricing page lists the mini realtime models at $10 per million audio input tokens and $20 per million audio output tokens, the full gpt-realtime-2 and 2.1 at $32 and $64, tts-1 at $15 per million characters and tts-1-hd at $30. For a finished voiceover, a character-priced model such as tts-1 is the simpler comparison. A realtime model is built for conversation, and its audio-token bill depends on how long it talks.
Sume's TTS Router prices Sonic voices at about $47.50 per million characters, which is higher than OpenAI's tts-1 list price. The difference buys a different engine, saved voices, avatar handles and word timestamps, not a lower bill.
Price table
OpenAI figures come from its pricing page. Sume's figure is computed from Cartesia list price times 1.25, rounded up to the cent for each job.
| Product | Unit | Price |
|---|---|---|
| gpt-realtime-2 / 2.1 | Per 1M audio tokens in / out | $32 / $64 |
| gpt-realtime mini | Per 1M audio tokens in / out | $10 / $20 |
| tts-1 | Per 1M characters | $15 |
| tts-1-hd | Per 1M characters | $30 |
| Sume TTS Router (Sonic) | Per 1M characters | About $47.50 |
Pick the right product in four steps
- Decide whether a person talks to it live (realtime) or you hand over a script (narration).
- For narration, compare per-character prices and a blind listening test.
- For realtime, measure audio tokens per minute from a real call before you estimate cost.
- Keep the voice and model id with every output so you can reproduce it.
import os, requests
r = requests.post(
"https://api.sume.com/v1/tts-router/generate",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "tts-demo-001",
},
json={
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare three prices.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"timestamps": {"words": True},
},
timeout=30,
)
r.raise_for_status()
print(r.json())Why the units matter
A character price is knowable before you generate: count the script. A token price needs a measured ratio. If you compare $20 per million output tokens with $47.50 per million characters you are comparing unlike things. Convert both to cost per finished minute using your own measured ratios, and keep the working in a sheet so you can update it when a price moves.
A worked narration example
A 1,000-character script costs $0.015 on tts-1, $0.030 on tts-1-hd and about 5 cents on Sume (4.75 cents rounded up to the cent for the job). For one script the gap is a few cents. For 10,000 scripts it is a few hundred dollars, which is real money but small next to the cost of a voice that does not suit your brand.
The right comparison therefore starts with a listening test. Render the same 1,000-character script on each engine, play them back to back to someone who does not know which is which, and only then bring price into the decision. If two voices are close, take the cheaper one. If one is clearly better, the price difference is usually the cheapest part of your video production.
What Sume does not do
Sume does not offer OpenAI voices, realtime conversation, or streaming audio. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. I did not verify how many audio tokens a minute of speech uses in OpenAI's models, so no conversion is given here.
Next
See realtime transcription priced per minute against file STT and the per-1,000-character ranking.
Sources
Related posts
More in Comparisons
- OpusClip Pro MCP connector vs the Sume hosted MCP: what differs
OpusClip lists a Video Editing API, Scheduler API and MCP connector on Pro. Sume has a hosted MCP with trim and inspect tools. What each gives an agent.
- Own agent loop vs one Sume Agent Completion request
POST /v1/agent/completions runs the Sume Agent with tools and media generation in one async call: required spend cap, 30 images, a schema and a receipt.
- Perfume ad shot at 1080p: H3 Max $2.00 vs Kling 3 $1.40 for 10 s
A 10-second 1080p product shot is $2.00 on MiniMax H3 Max and $1.40 on Kling 3 silent, $2.10 with sound. What the extra reference inputs on H3 Max buy.
- Phonon-2 (164 MB, English) or hosted Sume STT for a 10-minute file?
Phonon-2 is a free 164 MB English-only on-device model. Sume STT is hosted, multilingual, about 1 cent a minute. When each fits, with sourced numbers.
Written by Sume