A 1,300-second narration fails tts_duration_exceeded: split the script
Sume TTS fails a request whose audio runs past 1,200 seconds with tts_duration_exceeded and captures no credit. Split the script at sentence ends.

If Sume TTS would synthesize more than 1,200 seconds of audio for one request, the job fails with tts_duration_exceeded and no credit is captured. A 1,300-second narration is 100 seconds over, so it has to go out as at least two requests. The 20,000-character request limit is a separate cap: a script can be under it and still run past 20 minutes of audio if it is spoken slowly.
Both limits come from Sume's OpenAPI description of the TTS route, read on 2026-10-09. The catalog lists the price at $0.0475 per 1,000 characters.
The two limits at a glance
Length in characters and length in seconds are different measures. A request must satisfy both. Characters decide the price and the request cap. Seconds decide whether the synthesized result is accepted.
| Limit | Value | What happens past it |
|---|---|---|
| Transcript length | 1 to 20,000 characters | request rejected |
| Synthesized audio | 1,200 seconds | fails with tts_duration_exceeded, no credit captured |
| Maximum price of one request | $0.95 (20,000 characters) | not applicable |
| Join of parts afterward | timeline audio concat, 20 parts per job, $0.01 | not applicable |
Splitting a script safely
Do not cut at a fixed character count. Cut after a sentence, so each part starts and ends cleanly. A sentence splitter that accumulates sentences until a character budget is reached is enough. Pick a budget well under the 1,200-second limit: if your first test shows a 3,000-character part lasts 250 seconds, a budget of 12,000 characters is about 1,000 seconds, which leaves a margin.
The function below splits a script into parts below a character budget without breaking a sentence. A sentence longer than the budget becomes its own part.
import re
def split_script(text, budget):
sentences = re.split(r"(?<=[.!?])\s+", text.strip())
parts, current = [], ""
for s in sentences:
if current and len(current) + 1 + len(s) > budget:
parts.append(current)
current = s
else:
current = (current + " " + s).strip()
if current:
parts.append(current)
return parts
script = "One. Two is a bit longer. Three ends it."
for i, p in enumerate(split_script(script, 20), 1):
print(i, len(p), p)Joining the parts afterward
Each part returns its own audio file on a Sume media URL. To make one file, use the timeline audio concat operation with up to 20 parts per job at a flat $0.01, which joins in the sample domain with no gap. For a 1,300-second script split into two parts, the join costs one cent. The two TTS requests cost their characters at $0.0475 per thousand.
Because a failed over-length request captures no credit, a mistake here costs time and not money. Still, measure the first part's duration before you submit the rest, so you do not wait on a result that will fail.
- Use the first result's duration_seconds to calibrate the character budget.
- Keep the same voice and language on every part so the join sounds like one take.
- Use an Idempotency-Key per part, so a retried part is not billed twice.
A calibration run
Because the failure depends on seconds and not characters, the safest first step is a short calibration. Submit the first 2,000 characters of the script, read duration_seconds from the result, and compute seconds per character. If 2,000 characters came back as 150 seconds, that is 0.075 seconds per character, so a 1,200-second ceiling is about 16,000 characters, and a budget of 12,000 leaves a quarter of headroom.
The calibration costs 2 x $0.0475 = $0.095 and the audio is part of the final narration, so nothing is wasted. Voices and speed settings change the rate, so repeat the calibration whenever you change either.
Sources
Related posts
More in Developers
- 15 Ideogram 4.5 edits with 5 images each: $1.125 at medium on Sume
Ideogram 4.5 on Sume edits the first image and takes up to four more as references. Fifteen medium edits cost 15 x $0.075 = $1.125.
- 164 days from DALL-E 3 shutdown to GPT Image 1, then 39 more
OpenAI removed DALL-E 3 on May 12, 2026, retires GPT Image 1 on Oct 23, then three more image ids on Dec 1. Day counts and the Sume ids that replace them.
- 26 dialogue lines, a 20-part concat cap: two $0.01 jobs, then offsets
Timeline audio concat takes 1-20 parts. A 26-line dialogue needs two concat jobs, $0.02 total, plus about $0.17 of TTS. How to chain them and keep the offsets.
- A 28-minute talk transcribed on Sume: 3 detach ranges plus STT = $0.31
Audio detach caps output at 900 s and STT estimates cap at 10 minutes. Here is how a 28-minute talk splits into three ranges, and what it costs: $0.31.
Written by Sume