Convert audio API prices to dollars per finished minute (Python)
Per 1K characters, per million tokens, per song, per minute: a small Python converter that puts TTS and music prices on one scale with explicit assumptions.

To compare audio prices, convert each to dollars per finished minute using assumptions you state out loud: characters of narration per minute, minutes per song, and audio tokens per second. The vendors publish units that do not line up, and the missing conversion factor is yours to supply.
The script below does the arithmetic. The three assumptions are inputs, so change them and the table changes.
The units on the price pages
Prices as read on 2026-10-03. Each is a different unit, which is why a side-by-side table is misleading until converted.
| Product | Price | Unit |
|---|---|---|
| ElevenLabs v4 (72% discount until Oct 12) | $0.022 | per 1K characters |
| ElevenLabs v4 Turbo | $0.011 | per 1K characters |
| Gemini 3.8 Flash TTS audio output (through end of 2026) | $9.00 | per 1M tokens |
| ElevenLabs music | $0.15 | per minute |
| Lyria 3.5 | $0.08 | per generated song |
| Sume Music | $0.125 | per accepted generation |
The converter
CHARS_PER_MIN depends on speaking rate and language, so measure it from your own scripts. TOKENS_PER_SEC is the number of audio output tokens per second of speech, which you can read from a real response's usage data. SONG_MIN is the average length of the songs you keep. None of these three has a default from a vendor page, so the script has placeholder values you must replace.
# Replace these three with numbers measured on your own content.
CHARS_PER_MIN = 900
TOKENS_PER_SEC = 25
SONG_MIN = 2.0
def per_minute():
return {
"ElevenLabs v4 (discounted)": 0.022 * CHARS_PER_MIN / 1000,
"ElevenLabs v4 Turbo": 0.011 * CHARS_PER_MIN / 1000,
"Gemini 3.8 Flash TTS": 9.00 * TOKENS_PER_SEC * 60 / 1_000_000,
"ElevenLabs music": 0.15,
"Lyria 3.5": 0.08 / SONG_MIN,
"Sume Music": 0.125 / SONG_MIN,
}
if __name__ == "__main__":
for name, cost in sorted(per_minute().items(), key=lambda kv: kv[1]):
print(f"{name:<30} ${cost:.4f} per minute")Things the converter cannot fix
Treat the output as a planning figure, not a quote.
- The 900 characters per minute and 25 tokens per second above are placeholders. Do not publish a comparison built on them.
- ElevenLabs v4's price is marked as a discount until Oct 12. After that date the page, not this script, is the source for the new rate.
- A song price is per generation. If you keep one take in three, multiply the per-minute cost by three for the true cost of a finished minute.
- Quality is not in the table. A cheaper minute that needs two retries is not cheaper.
Where Sume's own price fits
Sume's music docs describe a fixed price per accepted generation and no duration field, so the cost of a minute depends only on how long a track you ask for in the prompt. For anything else you plan to call through Sume, read the live rate from the catalog instead of copying a number from a blog post, including this one.
Sources
Related posts
More in Pricing
- ElevenLabs API cost for a 1-minute video: v4, Turbo, music, SFX
ElevenLabs lists v4 at $0.022 per 1K characters, v4 Turbo at $0.011, music at $0.15 per minute and SFX at $0.12 per minute. A worked one-minute video.
- FLUX API per-image price: what 500 images cost at $0.024 to $0.048
BFL lists pay-as-you-go FLUX pricing from $0.024 to $0.048 per image with no subscription. Budget tables for 500 and 2,000 images, plus a retry allowance.
- FLUX video upscale cost: megapixel-second examples for a 10 s clip
BFL prices video upscaling at $0.07 (Precise) or $0.10 (Creative) per megapixel-second. Worked totals for 1080p, 2K and 4K output over 10 seconds.
- Gemini Omni Flash: 5,792 tokens a second is about $0.10 a second
Google bills Gemini Omni Flash video at $17.50 per million tokens. At 5,792 tokens a second for 720p, that is about $0.10 a second. The math, by clip length.
Written by Sume