Sume TTS: 20,000 characters or 1,200 seconds, which fails first?
At 15 characters a second, 20,000 characters is 1,333 seconds, so the 1,200-second audio cap trips first. Split near 15,000 characters per job.

For most scripts the 1,200-second audio limit fails first. Sume TTS accepts up to 20,000 characters per request, but synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured. At an assumed 15 characters a second, 20,000 characters runs 1,333 seconds, so a full-length request is refused. Plan for about 15,000 characters per job.
The two limits
The characters limit is checked on your text. The duration limit is checked on the finished audio, so it depends on the voice, the language and any speed setting. The OpenAPI reference states both: a 20,000 character maximum for transcript, and the 1,200 second failure with no credit capture.
Which one trips, by speaking rate
Your rate is a property of your script and voice; measure it from one completed job, then use the table.
| Characters per second (assumed) | 20,000 characters lasts | Characters that reach 1,200 s | First limit hit |
|---|---|---|---|
| 12 | 1,667 s | 14,400 | 1,200 s duration |
| 15 | 1,333 s | 18,000 | 1,200 s duration |
| 18 | 1,111 s | 21,600 | 20,000 characters |
Split with margin and price it
Cut at paragraph boundaries so each job stays near 15,000 characters, about 1,000 seconds at 15 characters a second. That leaves 200 seconds of room for slower stretches. A 60,000 character script becomes four jobs of 15,000 characters. Each costs 15 x $0.0475 = $0.7125, and the total is 4 x $0.7125 = $2.85, the same as 60 x $0.0475 (Sume API catalog, read 2026-10-09). Splitting costs nothing extra in TTS.
Join the parts
Re-assemble the takes with a Timeline audio concat job: up to 20 parts, one gapless file, $0.01 per job. Keep the takes as WAV when you will join or lip-sync them, because MP3 adds priming padding at every edge. The Timeline audio docs list the part limits and the refusal codes. If the final file drives a render, remember that a render is billed per started output minute.
Why the failure is cheap but not free
The documented failure for an over-long take captures no credit, so a refused 18,000 character job does not charge you $0.855. It does cost a round trip and the time you waited for synthesis, which is why you want the split decided before submission and not after the error.
A practical routine is: count characters, divide by your measured characters per second, and if the estimate exceeds roughly 1,000 seconds, split. Count the way the API does, with spaces and punctuation included; counting words undershoots. Submit the parts with distinct Idempotency-Key values so a retry of part three never duplicates part two. When all parts are back, check each duration before the Timeline concat so a silent truncation cannot hide inside a long file.
A practical rule
Pick the limit you will hit first from your speaking rate. Divide 20,000 characters by your measured characters per second: at 15 a second the character cap allows 1,333 seconds, so the 1,200 second duration limit arrives first and a request near the cap can fail with no credit captured. At 18 a second the cap allows 1,111 seconds and the character limit is the tighter one. Split on paragraph breaks, send each part as its own job with its own Idempotency-Key, and join the results with a Timeline audio concat for $0.01.
Sources
Related posts
More in Developers
- Sume TTS default is mp3: sentence slices need wav and emit_audio
Sume TTS defaults to mp3 at 44.1 kHz and 128 kbps. Per-sentence audio slices need wav (pcm_s16le) plus segmentation emit_audio, as in the Python request below.
- A 5-minute voice-over: 4.8 MB as 128 kbps MP3, 26.5 MB as WAV
Sume TTS defaults to MP3 at 44,100 Hz and 128 kbps. Five minutes is 4.8 MB; 16-bit mono WAV is 26.46 MB at 44.1 kHz and 28.8 MB at 48 kHz.
- TTS sentence slices: segmentation, wav output and boundary_lead_ms 70
Sume TTS can return gapless sentence segments. Needs timestamps.words and wav or raw for per-sentence audio_url. A 900-character script costs $0.04275.
- Sume TTS sentence segments: boundary_lead_ms 0 vs 500, 12 lines
Ask Sume TTS for words and sentence segmentation and you get gapless wav slices. boundary_lead_ms (default 70) moves each cut 0 to 500 ms after the last word.
Written by Sume