Read the TTS Router catalog in Python: price per 1M from list micros
Sume's TTS Router catalog publishes list micro-dollars per character; billing is list x 1.25. A Python script prints price per 1M and per 1,000 characters.

To get Sume's TTS price per 1M characters from code, read GET /v1/tts-router/models, take each row's list micro-dollars per character and multiply by 1.25. For the 38 micros per character list rate I read in Sume's notes that gives 47.5 micro-dollars, which is $47.50 per 1M characters, with the final charge rounded up to the cent. The catalog is the source to trust because it publishes the margin the reserve actually charges.
That matters this week because new models arrive with per-million prices. Microsoft lists MAI-Voice-2.1 at $22 and Flash at $15 per 1M characters (read 2026-10-03). To compare, put every price on the same unit, and let a script do it.
What the docs promise
The notes say the catalog exposes list micros per character, but I did not read the exact JSON field names, so the script below searches each row for a numeric field with 'micro' in its name instead of hard-coding one. If it finds none, it says so rather than guessing.
| Fact | Value |
|---|---|
| Catalog route | GET /v1/tts-router/models |
| Generate route | POST /v1/tts-router/generate |
| Billing basis | Cartesia list micros per character x 1.25, ceil to cents |
| Worked rate | 38 micros x 1.25 = 47.5 micros = $47.50 per 1M characters |
| Transcript length | 1 to 20,000 characters |
Unknown model | 400 with catalog_url |
| Streaming | Not offered |
The script
It fetches the catalog, finds the list micros per character on each row and prints price per 1,000 characters and per 1M characters at the 1.25 margin. A second helper prices a script of a given length, using the cent rounding the docs describe.
import math
import os
import requests
MARGIN = 1.25
r = requests.get(
"https://api.sume.com/v1/tts-router/models",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
timeout=30,
)
r.raise_for_status()
body = r.json()
rows = body if isinstance(body, list) else body.get("models") or body.get("data") or []
def micros(row):
for key, val in row.items():
if "micro" in key.lower() and isinstance(val, (int, float)):
return val
return None
def cost_usd(chars, list_micros):
return math.ceil(chars * list_micros * MARGIN / 10_000) / 100 # micros -> cents, rounded up
for row in rows:
m = micros(row)
name = row.get("id") or row.get("model")
if m is None:
print(name, "no micros field; read the row by hand")
continue
print(f"{name}: ${m * MARGIN / 1000:.4f} per 1k chars, ${m * MARGIN:.2f} per 1M, 900 chars = ${cost_usd(900, m):.2f}")
Turning price into a budget
A rate per million characters is hard to feel. Convert it to something your team plans in. At the 900 characters per minute planning figure (150 spoken words at about 6 characters each; adjust for your language), an hour of narration is about 54,000 characters, which is $2.57 at $47.50 per 1M, before the cent rounding on each job. A 20,000 character transcript, the per-request ceiling, costs $0.95.
Rounding up to the cent matters when you generate many short lines. A 40 character line costs a fraction of a cent in raw terms but is charged as at least a cent, so a script split into 120 one-sentence jobs costs more than the same text sent in a few large ones. If the join is free of audible seams, fewer requests is cheaper; if you need to retake single lines, more requests is cheaper to fix. Decide which matters more for the project.
When the script finds nothing
If the output says no micros field, do not guess. Open the raw response, copy one row, and read the price field by name. Then pin the field name in your script and add a test that fails when the field disappears, so a catalog change shows up in CI rather than in an invoice.
Also remember the catalog lists routable ids, and job.model echoes the id you asked for, so aliases such as sonic-latest keep working while the engine behind them changes. Log the id you sent and the price you computed next to each job so you can reconcile later.
Checks before you trust the output
Compare the result with the margin the catalog itself publishes. Notes in the docs say the older 1.10 figure is retired, so a script that hard-codes 1.10 will under-quote by about 12 percent. If a row's margin differs from 1.25, use the row's value.
Then run the same function on competitors' per-million prices so every number in the sheet is a rate per 1M characters, and keep a column for the date you read each one. Prices move, and a stale cell is the most common cause of a wrong estimate.
Sources
Related posts
More in Developers
- Read twelve narration takes at once: jobs_result partial success
Over MCP, jobs_result takes up to 20 job ids and returns one ok-or-error entry each. How to read a wave of TTS takes when one is still running.
- Recast request timed out: retry with the same Idempotency-Key
A timeout on POST /v1/video-router/generate does not tell you whether the job exists. Derive a stable Idempotency-Key so a retry returns the original job.
- Recast sync wait caps at 30 seconds: use a webhook for long sources
Sync and subscribe modes wait at most 30 seconds, so submit h3-max-recast jobs as async or webhook and poll /v1/jobs/{id}/status as a fallback.
- Recraft V4 on Sume returns WebP only: convert to PNG or JPEG in Python
Recraft V4 on Sume outputs WebP and takes no references. A Pillow converter for PNG or JPEG, with transparency flattened onto white for JPEG delivery.
Written by Sume