Convert audio API prices to dollars per finished minute (Python)

Per 1K characters, per million tokens, per song, per minute: a small Python converter that puts TTS and music prices on one scale with explicit assumptions.

4 min readSume
All posts

To compare audio prices, convert each to dollars per finished minute using assumptions you state out loud: characters of narration per minute, minutes per song, and audio tokens per second. The vendors publish units that do not line up, and the missing conversion factor is yours to supply.

The script below does the arithmetic. The three assumptions are inputs, so change them and the table changes.

The units on the price pages

Prices as read on 2026-10-03. Each is a different unit, which is why a side-by-side table is misleading until converted.

Audio list prices by unit (read 2026-10-03)
ProductPriceUnit
ElevenLabs v4 (72% discount until Oct 12)$0.022per 1K characters
ElevenLabs v4 Turbo$0.011per 1K characters
Gemini 3.8 Flash TTS audio output (through end of 2026)$9.00per 1M tokens
ElevenLabs music$0.15per minute
Lyria 3.5$0.08per generated song
Sume Music$0.125per accepted generation

The converter

CHARS_PER_MIN depends on speaking rate and language, so measure it from your own scripts. TOKENS_PER_SEC is the number of audio output tokens per second of speech, which you can read from a real response's usage data. SONG_MIN is the average length of the songs you keep. None of these three has a default from a vendor page, so the script has placeholder values you must replace.

# Replace these three with numbers measured on your own content.
CHARS_PER_MIN = 900
TOKENS_PER_SEC = 25
SONG_MIN = 2.0

def per_minute():
    return {
        "ElevenLabs v4 (discounted)": 0.022 * CHARS_PER_MIN / 1000,
        "ElevenLabs v4 Turbo": 0.011 * CHARS_PER_MIN / 1000,
        "Gemini 3.8 Flash TTS": 9.00 * TOKENS_PER_SEC * 60 / 1_000_000,
        "ElevenLabs music": 0.15,
        "Lyria 3.5": 0.08 / SONG_MIN,
        "Sume Music": 0.125 / SONG_MIN,
    }

if __name__ == "__main__":
    for name, cost in sorted(per_minute().items(), key=lambda kv: kv[1]):
        print(f"{name:<30} ${cost:.4f} per minute")

Things the converter cannot fix

Treat the output as a planning figure, not a quote.

  • The 900 characters per minute and 25 tokens per second above are placeholders. Do not publish a comparison built on them.
  • ElevenLabs v4's price is marked as a discount until Oct 12. After that date the page, not this script, is the source for the new rate.
  • A song price is per generation. If you keep one take in three, multiply the per-minute cost by three for the true cost of a finished minute.
  • Quality is not in the table. A cheaper minute that needs two retries is not cheaper.

Where Sume's own price fits

Sume's music docs describe a fixed price per accepted generation and no duration field, so the cost of a minute depends only on how long a track you ask for in the prompt. For anything else you plan to call through Sume, read the live rate from the catalog instead of copying a number from a blog post, including this one.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume