ElevenLabs v4 vs Gemini TTS: cost per hour with the assumptions shown
One hour of narration priced from reported rates: ElevenLabs v4 at $0.08 per 1,000 characters, Gemini TTS at 25 audio tokens per second. Python included.

One hour of narration costs about $4.32 on ElevenLabs v4 at its listed $0.08 per 1,000 characters, assuming 54,000 characters per hour, and about $0.81 on Gemini Flash TTS output, assuming 25 audio tokens per second and $9.00 per million output tokens. Every number depends on assumptions you can change, so the script below shows them.
The rates come from Digital Applied's September tracker and the Gemini API changelog, both read 2026-10-02. The per-million-token unit for Gemini is our reading of the tracker, and the characters-per-hour figure is our assumption (about 150 words a minute at six characters per word).
What are the assumptions?
Change any row and the answer changes.
| Provider and model | Rate used | Hour of speech |
|---|---|---|
| ElevenLabs v4 | $0.08 per 1,000 characters | $4.32 |
| ElevenLabs Turbo | $0.04 per 1,000 characters | $2.16 |
| ElevenLabs v4, promotion to 2026-10-12 | $0.022 per 1,000 characters | $1.19 |
| Gemini Flash TTS output | $9.00 per 1M tokens, 90,000 tokens | $0.81 |
| Gemini Flash-Lite TTS output | $6.00 per 1M tokens, 90,000 tokens | $0.54 |
Why is the gap so large, and does it mean anything?
The two bills count different things: characters in, versus audio tokens out. Price per hour ignores voice quality, languages, voice control and limits, and the Gemini rates are promotional through 2026-12-31. Compare on a sample you listen to, not on a table.
chars_per_hour = 54_000 # assumption: 150 wpm x 6 chars
for name, per_1k in [('v4', 0.08), ('turbo', 0.04), ('v4 promo', 0.022)]:
print(name, round(chars_per_hour / 1000 * per_1k, 2))
tokens_per_hour = 25 * 3600 # 25 audio tokens per second
for name, per_m in [('gemini flash', 9.00), ('gemini flash-lite', 6.00)]:
print(name, round(tokens_per_hour / 1e6 * per_m, 2))Where does Sume fit in a comparison like this?
Sume's TTS 1.0 lists a transcript limit of 20,000 characters per job and fails audio longer than 1,200 seconds with tts_duration_exceeded. The docs do not publish a Sume dollar rate for it, so there is no row here; run a short job and read the billed amount.
Sources
Related posts
More in Comparisons
- ElevenLabs stability 0.5 and similarity 0.75 vs Sume generation_config
ElevenLabs defaults stability to 0.5 and similarity to 0.75. Sume has no such sliders: generation_config takes only volume, speed and an emotion string.
- fal Agent (early access) vs Sume Agent Completions and hosted MCP
fal's August 2026 agent is early access, with API, CLI and MCP. Sume's agent has Agent Completions over the API and a hosted MCP. What each documents today.
- fal cancel returns 202 or 400: how Sume's cancel 409 differs
fal's queue cancel answers 202 CANCELLATION_REQUESTED, 400 ALREADY_COMPLETED or 404. Sume's cancel works only before generation starts, else 409.
- fal media lasts 7 days by default; Sume URLs do not expire
fal keeps generated media on its CDN for at least 7 days and lets you set a lifecycle per request. Sume media.sume.com URLs do not expire. What to store.
Written by Sume