Estimate a Sume TTS bill in Python before you submit: count characters
Sume TTS is $0.0475 per 1,000 characters, spaces and punctuation included. A small Python function prices a script and flags ones near the limits.

The formula
The cost of a Sume TTS job is the number of characters in the transcript times $0.0475 divided by 1,000. A 5,000-character script is 5 x $0.0475 = $0.2375. Spaces and punctuation count, so measure the exact string you will send.
A request can carry 1 to 20,000 characters, but a job also fails with tts_duration_exceeded if the audio runs past 1,200 seconds. That second limit usually bites first.
Examples
Use the table to sanity-check your function.
| Characters | Cost | Notes |
|---|---|---|
| 1,000 | $0.0475 | About a minute and a quarter of speech |
| 5,000 | $0.2375 | A short video script |
| 15,000 | $0.7125 | Close to the 1,200-second cap |
| 20,000 | $0.95 | The request character limit |
The function
The code uses Decimal so the arithmetic does not pick up float noise. It warns when a script is likely to exceed 1,200 seconds, using Cartesia's published rule of thumb that a minute of audio is about 750 to 800 characters.
from decimal import Decimal
RATE = Decimal("0.0475") # USD per 1,000 characters
MAX_CHARS = 20000
CHARS_PER_MIN_LOW = 750 # Cartesia pricing page rule of thumb
def estimate(text):
n = len(text)
if not 1 <= n <= MAX_CHARS:
raise ValueError(f"{n} characters is outside 1-{MAX_CHARS}")
cost = (Decimal(n) * RATE / 1000).quantize(Decimal("0.0001"))
minutes = n / CHARS_PER_MIN_LOW
return cost, minutes
if __name__ == "__main__":
script = "Welcome to the Q4 review. " * 200
cost, minutes = estimate(script)
print(f"${cost} for about {minutes:.1f} minutes at the slow end")
if minutes > 20:
print("Split this script: a job over 1,200 seconds fails")Using it well
Cartesia's pricing page says one minute of audio is about 750 to 800 credits, with a credit being one character. The rule is for planning only. Real timing depends on voice, speed and punctuation. If your estimate is near 15,000 characters, split the script into chapters or scenes and submit each separately. The same code then tells you the cost of every part, and you can store it next to the job id.
Python's len counts code points, which is what most scripts are made of. If you use emoji or combining marks heavily, test one sample against the job's billed usage before you rely on the estimate.
- Strip stage directions you do not want spoken before you count.
- Estimate per scene, not just per script, to find the expensive part.
- Log the estimate and compare it with the final invoice once a month.
- Prices change; read the current rate from
GET /v1/catalogif you automate budgets.
Putting it in a pipeline
Call the estimate before every submit and refuse to send a script that would cost more than a limit you set, for example $2 for a single job. That keeps a pasted novel from becoming a surprise charge. Log the job id, the character count and the estimate in the same line, so a later invoice check is one search.
For a batch of slides or announcements, estimate each item and sum them. A deck of 20 slides at 600 characters is $0.57, and the code should print exactly that. If the printed value differs from your hand arithmetic, fix the code before it touches a paid route.
Finally, remember that the estimate is a ceiling for planning. Billing follows the characters the API receives, and a request that fails validation before it runs does not create audio. Retries with the same Idempotency-Key do not create a second charge, which is why you should always send one.
Sources
Related posts
More in Developers
- TTS 400: Provide exactly one of transcript_source or transcript
This 400 means the TTS body had both transcript and transcript_source, or neither. Send exactly one, plus a voice, and re-run.
- TTS 400 asking for voice.id or avatar_id / avatar_handle: the fix
The TTS call needs a voice: set voice.id, or a top-level avatar_id or avatar_handle. Without one the request is rejected before any audio is made.
- List Sume TTS Router models before you hardcode a Sonic id
GET /v1/tts-router/models returns the catalog of pass-through TTS models. Read it at startup instead of pinning an id that may change.
- Voice replication API audit checklist before you switch
Gemini 3.8 Flash TTS is GA with voice replication and 150+ voices. Before switching providers, audit these items against Sume's live catalog.
Written by Sume