How many minutes of speech fit in one 20,000-character TTS request?
Sume TTS caps a request at 20,000 characters. At 800 to 1,200 characters per minute that is roughly 17 to 25 minutes of audio and costs 95 cents.

The limit and what it means in minutes
Every Sume TTS Router model accepts up to 20,000 characters per request. The number of minutes that fits depends on how fast the voice speaks your text, which varies by language, punctuation and the speed setting. Use the table as a planning range, then measure one real take.
A full 20,000-character request costs 95 cents on Sume. Sume TTS list price is 38 micro-dollars (0.0038 cents) per character; the billable price is list x 1.25, rounded up to whole cents per job. That is 0.00475 cents per character before rounding, so a 20,000-character request is 20,000 x 0.00475 = 95 cents.
Planning table
The minutes below are simple division of 20,000 by an assumed characters-per-minute rate. They are assumptions for planning, not measured outputs.
| Assumed chars per minute | Minutes in 20,000 chars | Sume cost per full request | Sume cost per minute |
|---|---|---|---|
| 800 | 25.0 | 95 cents | 3.80 cents |
| 900 | 22.2 | 95 cents | 4.28 cents |
| 1000 | 20.0 | 95 cents | 4.75 cents |
| 1200 | 16.7 | 95 cents | 5.70 cents |
When the text is longer
Split on paragraph or sentence boundaries, never mid-sentence, send each part as its own job, and join the results. Joining is a separate Sume job that concatenates Sume-hosted audio sample-exactly with no re-synthesis (Timeline audio), at $0.01 per job and up to 1,800 seconds of output.
If the joined narration runs past 30 minutes, the 1,800-second cap means you need two joined files. Keep a log of the character count per part so you can predict minutes before you pay.
Keeping parts consistent
Read the finished job to see the model, voice, language and speed it used, and send the same values for the next part (Jobs and results). For repeatable output across parts, pin a concrete model id rather than relying on an alias that moves.
Measure one take first
Before you split a long script, generate a 1,000-character sample at the voice and speed you plan to use, and note the audio duration. Multiply to get a characters-per-minute figure for your text. That costs 1,000 x 0.00475 = 4.75 cents, rounded up to 5 cents, and saves guessing.
Speed settings change the result. A slower speed means fewer characters per minute and therefore more minutes in a full request. Record the speed you used next to the figure so future scripts use the same basis.
Why not always send the maximum
A full request is the cheapest per character only through rounding, and a single failure wastes the whole job. Sending 5,000 to 10,000 characters per job keeps retries cheap: a failed 5,000-character job costs at most 5,000 x 0.00475 = 23.75 cents, rounded up to 24 cents, while a retried 20,000-character job costs 95 cents. Pick part sizes that match how you review the script, one chapter or scene per job.
Sources
Related posts
More in Developers
- Hy Image 3.5 Preview returns base64 PNG; Sume returns a URL
Moving image code from a base64 PNG response (Hy Image 3.5 Preview on OpenRouter) to Sume means downloading data[].url. A 12-line Python version.
- One idempotency key per prompt row: re-run a 1,000-clip library safely
Derive the Idempotency-Key from every field you send, so a crashed batch can restart without paying twice and a changed row never collides. Python, 17 lines.
- Image 1.0 defaults to low quality; GPT Image 2.5 defaults to high
Same prompt, two Sume routes, two defaults: Image 1.0 omits quality as low, POST /v1/images with GPT Image 2.5 omits it as high. Set quality in both.
- Image 1.0 to POST /v1/images: image_urls becomes input_references
Move a Sume Image 1.0 request to POST /v1/images: image_urls to input_references, mask_image_url to mask_url, num_images to n, plus new defaults.
Written by Sume