Sume TTS: 20,000 characters or 1,200 seconds, which fails first?

At 15 characters a second, 20,000 characters is 1,333 seconds, so the 1,200-second audio cap trips first. Split near 15,000 characters per job.

5 min readSume
All posts

For most scripts the 1,200-second audio limit fails first. Sume TTS accepts up to 20,000 characters per request, but synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured. At an assumed 15 characters a second, 20,000 characters runs 1,333 seconds, so a full-length request is refused. Plan for about 15,000 characters per job.

The two limits

The characters limit is checked on your text. The duration limit is checked on the finished audio, so it depends on the voice, the language and any speed setting. The OpenAPI reference states both: a 20,000 character maximum for transcript, and the 1,200 second failure with no credit capture.

Which one trips, by speaking rate

Your rate is a property of your script and voice; measure it from one completed job, then use the table.

Which Sume TTS limit trips first at three assumed speaking rates; limits from the Sume OpenAPI reference, read 2026-10-09.
Characters per second (assumed)20,000 characters lastsCharacters that reach 1,200 sFirst limit hit
121,667 s14,4001,200 s duration
151,333 s18,0001,200 s duration
181,111 s21,60020,000 characters

Split with margin and price it

Cut at paragraph boundaries so each job stays near 15,000 characters, about 1,000 seconds at 15 characters a second. That leaves 200 seconds of room for slower stretches. A 60,000 character script becomes four jobs of 15,000 characters. Each costs 15 x $0.0475 = $0.7125, and the total is 4 x $0.7125 = $2.85, the same as 60 x $0.0475 (Sume API catalog, read 2026-10-09). Splitting costs nothing extra in TTS.

Join the parts

Re-assemble the takes with a Timeline audio concat job: up to 20 parts, one gapless file, $0.01 per job. Keep the takes as WAV when you will join or lip-sync them, because MP3 adds priming padding at every edge. The Timeline audio docs list the part limits and the refusal codes. If the final file drives a render, remember that a render is billed per started output minute.

Why the failure is cheap but not free

The documented failure for an over-long take captures no credit, so a refused 18,000 character job does not charge you $0.855. It does cost a round trip and the time you waited for synthesis, which is why you want the split decided before submission and not after the error.

A practical routine is: count characters, divide by your measured characters per second, and if the estimate exceeds roughly 1,000 seconds, split. Count the way the API does, with spaces and punctuation included; counting words undershoots. Submit the parts with distinct Idempotency-Key values so a retry of part three never duplicates part two. When all parts are back, check each duration before the Timeline concat so a silent truncation cannot hide inside a long file.

A practical rule

Pick the limit you will hit first from your speaking rate. Divide 20,000 characters by your measured characters per second: at 15 a second the character cap allows 1,333 seconds, so the 1,200 second duration limit arrives first and a request near the cap can fail with no credit captured. At 18 a second the cap allows 1,111 seconds and the character limit is the tighter one. Split on paragraph breaks, send each part as its own job with its own Idempotency-Key, and join the results with a Timeline audio concat for $0.01.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume