A 1,300-second narration fails tts_duration_exceeded: split the script

Sume TTS fails a request whose audio runs past 1,200 seconds with tts_duration_exceeded and captures no credit. Split the script at sentence ends.

3 min readSume
All posts

If Sume TTS would synthesize more than 1,200 seconds of audio for one request, the job fails with tts_duration_exceeded and no credit is captured. A 1,300-second narration is 100 seconds over, so it has to go out as at least two requests. The 20,000-character request limit is a separate cap: a script can be under it and still run past 20 minutes of audio if it is spoken slowly.

Both limits come from Sume's OpenAPI description of the TTS route, read on 2026-10-09. The catalog lists the price at $0.0475 per 1,000 characters.

The two limits at a glance

Length in characters and length in seconds are different measures. A request must satisfy both. Characters decide the price and the request cap. Seconds decide whether the synthesized result is accepted.

TTS request limits as of 2026-10-09 (OpenAPI and catalog read 2026-10-09)
LimitValueWhat happens past it
Transcript length1 to 20,000 charactersrequest rejected
Synthesized audio1,200 secondsfails with tts_duration_exceeded, no credit captured
Maximum price of one request$0.95 (20,000 characters)not applicable
Join of parts afterwardtimeline audio concat, 20 parts per job, $0.01not applicable

Splitting a script safely

Do not cut at a fixed character count. Cut after a sentence, so each part starts and ends cleanly. A sentence splitter that accumulates sentences until a character budget is reached is enough. Pick a budget well under the 1,200-second limit: if your first test shows a 3,000-character part lasts 250 seconds, a budget of 12,000 characters is about 1,000 seconds, which leaves a margin.

The function below splits a script into parts below a character budget without breaking a sentence. A sentence longer than the budget becomes its own part.

import re

def split_script(text, budget):
    sentences = re.split(r"(?<=[.!?])\s+", text.strip())
    parts, current = [], ""
    for s in sentences:
        if current and len(current) + 1 + len(s) > budget:
            parts.append(current)
            current = s
        else:
            current = (current + " " + s).strip()
    if current:
        parts.append(current)
    return parts

script = "One. Two is a bit longer. Three ends it."
for i, p in enumerate(split_script(script, 20), 1):
    print(i, len(p), p)

Joining the parts afterward

Each part returns its own audio file on a Sume media URL. To make one file, use the timeline audio concat operation with up to 20 parts per job at a flat $0.01, which joins in the sample domain with no gap. For a 1,300-second script split into two parts, the join costs one cent. The two TTS requests cost their characters at $0.0475 per thousand.

Because a failed over-length request captures no credit, a mistake here costs time and not money. Still, measure the first part's duration before you submit the rest, so you do not wait on a result that will fail.

  • Use the first result's duration_seconds to calibrate the character budget.
  • Keep the same voice and language on every part so the join sounds like one take.
  • Use an Idempotency-Key per part, so a retried part is not billed twice.

A calibration run

Because the failure depends on seconds and not characters, the safest first step is a short calibration. Submit the first 2,000 characters of the script, read duration_seconds from the result, and compute seconds per character. If 2,000 characters came back as 150 seconds, that is 0.075 seconds per character, so a 1,200-second ceiling is about 16,000 characters, and a budget of 12,000 leaves a quarter of headroom.

The calibration costs 2 x $0.0475 = $0.095 and the audio is part of the final narration, so nothing is wasted. Voices and speed settings change the rate, so repeat the calibration whenever you change either.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume