How many characters is one minute of TTS audio? 1,200 s math
A vendor rule of thumb says a minute of speech is 750-800 characters. Here is what that means for a 1,200-second Sume TTS job and its 20,000-char cap.

Why the character count and the clock both matter
Sume TTS 1.0 limits a request two ways. The transcript can hold at most 20,000 characters, and the finished audio must not run past 1,200 seconds. If the audio is longer than 1,200 seconds the job fails with tts_duration_exceeded and no credits are captured. So before you submit a long script you want to know which limit you will hit first.
Cartesia's pricing page gives a planning figure: it says one minute of audio generation requires 750-800 credits, and that one credit equals one character. That is the vendor's own statement about its engine, so treat it as an estimate for a normal speaking pace rather than a guarantee for your voice and your speed setting.
The arithmetic
At 750 to 800 characters per minute, a 1,200-second job is 20 minutes of audio. That is 15,000 characters at the low rate and 16,000 at the high rate. Both are below the 20,000-character cap, so on a normal pace the duration cap, not the character cap, is the one that stops you.
Turn it around: 20,000 characters at the same rates would be roughly 25 to 26.7 minutes of audio. A script that long will probably cross 1,200 seconds and fail. Split it at a paragraph or chapter break instead of hoping it fits.
| Audio length | Characters at 750/min | Characters at 800/min |
|---|---|---|
| 1 minute | 750 | 800 |
| 5 minutes | 3,750 | 4,000 |
| 10 minutes | 7,500 | 8,000 |
| 20 minutes (1,200 s cap) | 15,000 | 16,000 |
What a safe split looks like
Pick a target well under the cap. A 12,000-character chunk is about 15 to 16 minutes at the vendor rate, which leaves room for a slower voice or a lowered speed. Slowing speech to 0.6 stretches the same text toward a much longer clip, so a setting that makes the voice slower pushes you toward the 1,200-second wall sooner.
- Split on sentence or paragraph boundaries so each file starts and ends cleanly.
- Keep each chunk below roughly 15,000 characters unless you have measured your own voice.
- Submit each chunk with its own Idempotency-Key so a retry returns the original job instead of paying twice.
- Stitch the finished files with Timeline audio if you need one file.
Checking your own voice
The only number that counts for your account is a measured one. Run one representative chunk, read the duration of the returned audio, and divide the character count by the minutes. If you see 650 characters per minute, your 1,200-second ceiling is about 13,000 characters, not 15,000.
The price side is simple: the Sume list price is $0.0475 per 1,000 characters, so 15,000 characters is $0.7125 and 20,000 is $0.95. See text-to-speech API cost per minute for the per-minute view, and the API pricing page for the current book.
Worked example: a 45-minute audiobook chapter set
Say a book section runs 36,000 characters. At 750 to 800 characters per minute that is about 45 to 48 minutes of speech, so it needs at least three jobs, because each job is capped at 1,200 seconds. Split into three pieces of 12,000 characters each and every piece lands at about 15 to 16 minutes, with headroom.
The cost does not change with the split: 36,000 characters at $0.0475 per 1,000 is $1.71. What changes is the retry surface. A failure in piece two costs you piece two, not the whole book.
- Count characters before you call, using the same rule the API uses.
- Name each piece with its chapter and part so the stitch order is obvious.
- Check the first piece's real duration, then adjust the other two.
Takeaway
Plan chunks of about 12,000 to 15,000 characters, measure one, and let tts_duration_exceeded be the signal you set the chunk size too high. No credits are captured on that failure, but a split plan avoids the wasted wait.
Sources
Related posts
More in Developers
- Idempotency-Key over 255 characters: hash long business keys
Sume accepts Idempotency-Key values up to 255 characters. Keep readable keys when short and fall back to a prefixed SHA-256 for long ones. Runnable Python.
- Idempotency-Key for a SaaS: customer, order and version
Derive a Format run's Idempotency-Key from customer id, order id, Format slug and a version you bump on purpose, so double clicks never make a second paid run.
- Ideogram 4 download: Hugging Face gate, login and first image
To run Ideogram 4 locally: accept the gate on Hugging Face, log in with hf, pip install the repo, run run_inference.py. The flags and the nf4 or fp8 choice.
- Image batch in Python: read ratelimit headers and retry-after
Sume can send ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after on image calls. Back off on 429 in Python, and treat queue_full separately.
Written by Sume