Split a long script into Sume TTS requests and price it
Sume TTS takes 20,000 characters per request at $0.0475 per 1,000. A Python splitter cuts at sentence ends and prices each request: 99,000 characters is $4.70.

A script of 99,000 characters needs 5 Sume TTS requests and costs about $4.70 at $0.0475 per 1,000 characters, because each request takes at most 20,000 characters. The Python below splits a long script at sentence ends and prices every request, so you know the count and the bill before you submit anything.
What the limits are
Sume TTS 1.0 is priced at $0.0475 per 1,000 characters on the API pricing page, which is the provider list of $0.038 times 1.25, and a single request takes up to 20,000 characters. A full request therefore costs $0.95. Spaces and punctuation count as characters, so the length of the string you send is the number that is billed.
The price per character is the same whether a request is short or long, so splitting costs nothing extra. The reason to split cleanly is audio quality at the seams, not money.
The splitter and pricer
The function looks for the last sentence end (a period followed by a space) before the 20,000-character limit and cuts there. If a sentence is longer than the limit, it cuts at the limit. The pricing function works in micros, rounds each request up, and sums them. The sample script is one line of narration repeated 2,200 times, so you can see the numbers without bringing your own text.
import math
MICROS_PER_CHAR = 47.5 # $0.0475 per 1,000 characters
MAX_CHARS = 20000 # one TTS request
def split(text, limit=MAX_CHARS):
chunks = []
while len(text) > limit:
cut = text.rfind(". ", 0, limit)
cut = cut + 1 if cut > 0 else limit
chunks.append(text[:cut])
text = text[cut:].lstrip()
if text:
chunks.append(text)
return chunks
def cost_usd(chunks):
micros = sum(math.ceil(len(c) * MICROS_PER_CHAR) for c in chunks)
return micros / 1_000_000
script = "A short line of narration for the voiceover. " * 2200
chunks = split(script)
print(len(script), "characters")
print(len(chunks), "requests")
print(max(len(c) for c in chunks), "characters in the largest")
print(f"${cost_usd(chunks):.4f} total")What it prints
Running the code on 2026-10-05 printed these four lines.
- 99000 characters
- 5 requests
- 19979 characters in the largest
- $4.7023 total
Reading the output
The split is by sentence, so each request lands a little under the limit rather than exactly on it: the largest is 19,979 characters. Five requests cover 99,000 characters because 99,000 divided by 20,000 is 4.95. The bill is linear in characters: 99,000 times 47.5 micros would be $4.7025, and the script prints $4.7023 because the spaces dropped at each cut are not sent, so slightly fewer characters are billed.
If you want fewer requests, raise the limit to the maximum and accept cuts that land mid-sentence. If you want smooth audio, keep the sentence rule and accept that the largest request sits a few hundred characters under 20,000.
Check the seams
Two requests that meet mid-paragraph may differ slightly in pacing, since each request is a separate generation. The simplest fix is to cut at paragraph breaks when you can and use the sentence rule only as a fallback. Another is to listen to the join once per script, since a seam problem shows up in the first listen and costs only a re-record of the shorter neighbouring request.
None of this changes the price. If the script has headings, speaker labels or stage directions that you do not want read aloud, strip them before you count, because every character you send is billed whether or not you meant it to be spoken. A quick pass that removes markup can cut a few percent from a long script. A request of 1,000 characters is $0.0475, a request of 20,000 is $0.95, and the total for the same text is the same however you cut it, to within the micro rounding.
Using it for a real script
Replace the sample text with your script, keep the limit at 20,000 and check the printed total against the usage dashboard after your first run. Keep one request per section, and name each one after the section so the usage line is easy to match, such as chapters or scenes, if you will want to re-record a part without paying for the whole script again. A re-recorded 3,000-character scene costs $0.1425 and not the $4.70 of a full re-run.
Sources
Related posts
More in Pricing
- Twelve 750-character match recaps a weekend: a year of TTS cost
Twelve recaps of 750 characters every weekend for 52 weeks is 468,000 characters: $7.02 on MAI Flash, $10.30 on MAI-Voice-2.1, $22.23 on Sume TTS.
- Station announcements: 50 messages at 160 characters, Flash vs Sume
Fifty pre-recorded announcements of 160 characters are 8,000 characters: 12 cents on MAI-Voice-2.1-Flash, 18 cents on MAI-Voice-2.1 and 38 cents on Sume.
- Stocking-stuffer gift videos: a bigger spend cap for the hero SKU only
In one bulk queue every item can carry its own generation_spend_cap_usd. Give the hero gift a higher ceiling and cheap stocking stuffers a low one.
- STT duration_seconds omitted: Sume reserves one minute. What to send
When you leave duration_seconds out of a Sume transcribe request, one audio minute is reserved. Send the real length, up to 600 seconds, to size the hold.
Written by Sume