Cartesia plan price per 1,000 characters vs Sume TTS on Sonic
Cartesia's plans work out to $0.0374 to $0.05 per 1,000 characters; Sume's pay-as-you-go Sonic rate is $0.0475. Where each is cheaper, with the arithmetic.

The comparison
Per 1,000 characters, Cartesia's Pro plan works out to $0.05, Startup to about $0.0392 and Scale to about $0.0374, while Sume's pay-as-you-go TTS is $0.0475. Sume is cheaper than Pro and cheaper than the other two plans if you use less than the plan's credits.
Cartesia counts one credit per character, and the Sonic models are the same engine family that Sume's POST /v1/tts-router/generate calls (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview).
Plan arithmetic
The plan figures come from Cartesia's pricing page. Divide the price by the credits, then by 1,000. Free is 20,000 credits and costs nothing.
| Plan | Monthly price | Credits | Per 1,000 characters | Sume buys the same dollars as |
|---|---|---|---|---|
| Pro | $5 | 100,000 | $0.0500 | 105,263 characters |
| Startup | $49 | 1,250,000 | $0.0392 | 1,031,579 characters |
| Scale | $299 | 8,000,000 | $0.0374 | 6,294,737 characters |
Where each wins
Read the last column as a break-even. If you use fewer characters in a month than that number, Sume costs less than the plan, because you pay only for characters you send. If you use more, the plan is cheaper per character. For Scale that crossover is about 6.3 million characters, and for Startup about 1.03 million.
That ignores unused credits and any overage rules, so read Cartesia's terms before you rely on the break-even. It also ignores everything that is not price: Sume gives you one jobs API with idempotency keys and webhooks shared with its other tools, plus pronunciation dictionaries and sentence segments from the same request.
Picking a model id
Send model on the router; it is required, and the rejected-on-1.0 rule means sume/tts-1.0 takes no model at all. Cartesia's model docs (read 2026-10-04) describe Sonic 3.6 with 44 languages. Use sonic-latest only if you accept silent upgrades, and pin sonic-3.6 for a series that must sound the same.
GET /v1/tts-router/modelslists what your key can call.- Set
languagefor every non-English transcript. - Characters include spaces and punctuation on both sides.
- Check Cartesia's current terms before you commit to a plan.
Turning minutes into characters
Cartesia's pricing page says one minute of audio is about 750 to 800 credits, and a credit is a character. So a 10-minute narration is roughly 7,500 to 8,000 characters. On Sume that is 7.5 x $0.0475 = $0.356 to 8 x $0.0475 = $0.38 per 10 minutes, and a plan only pays off when you generate hours of audio each month.
Sume's own page on how long a TTS minute is explains the same conversion from the 1,200-second job cap. Count your real script rather than guessing: send the exact text you will bill, since spaces and punctuation are characters on both services. A script with long numbers, URLs or stage directions is longer than it looks, so trim those before you compare plans, and keep one spreadsheet row per episode so the monthly total is a sum you can check against the invoice.
Sources
Related posts
More in Comparisons
- ChatGPT Image 2 vs 2.5 on Sume: $0.264 vs $0.066 per image
Sume's catalog shows ChatGPT Image 2 at $0.26375 and ChatGPT Image 2.5 at $0.065875, with 16 references, a mask_url and extra quality tiers on 2.5.
- Claude batch results wait for all; Sume webhooks fire per video
Anthropic Message Batches give results once every request has finished (or after 24 hours). Sume can notify per finished run. Which fits a holiday launch.
- Claude routines vs Desktop tasks vs /loop for recurring Sume renders
Pick between a Claude cloud routine, a Desktop task and /loop for scheduled Sume generation, or a Sume cron schedule when the work is only generation.
- Code2Video: 168 briefs and how to score your own
HeyGen's Code2Video benchmark uses 168 briefs from real launch videos. Here is how to build a smaller scoring set and render it from structured input on Sume.
Written by Sume