tts_duration_exceeded: split long scripts for Sume TTS
Sume TTS fails with tts_duration_exceeded when audio would pass 1,200 seconds, and takes at most 20,000 characters. Split rules and how to join the parts.

Sume TTS stops a job with tts_duration_exceeded when the synthesized audio would run past 1,200 seconds, which is 20 minutes, and no credit is captured for the failed job. A request also cannot carry more than 20,000 characters in transcript. Split long scripts at paragraph boundaries, generate each part, and join the takes.
The two caps
| Cap | Value | Rough equivalent |
|---|---|---|
| transcript length | 20,000 characters | about 3,300 words |
| audio length | 1,200 seconds | about 3,000 words at 150 words per minute |
| billing | $0.0475 per 1,000 characters, plus a 5.5% agent fee by default | 20,000 characters = $0.95 before the fee |
Why you may hit duration first
A slow voice or generation_config.speed of 0.6 stretches speech, so a script well under 20,000 characters can still pass 1,200 seconds. Speed accepts 0.6 to 1.5. If you slow the delivery, shrink the chunk size too.
A splitting recipe
- Split on blank lines, never mid-sentence.
- Pack paragraphs until a chunk reaches about 6,000 characters, a comfortable margin below both caps.
- Use the same model, voice and
generation_configfor every chunk, or the seams will be audible. - Number the files, submit chunks in parallel as async jobs, and wait for all of them with one
jobs_waitcall (or poll each job's status). - Join with Timeline audio concat (
POST /v1/timeline-1.0/audio,operation: concat), which joins in the sample domain with no added silence at the seams. It takes 1 to 20 parts per job, so a very long script needs a second pass.
Seam quality
Each chunk starts with fresh intonation, so end chunks at a natural pause, like a heading or a full stop at the end of a paragraph. Listen to each joint once. segmentation on TTS with timestamps.words can give you sentence timing if you need captions, with the slice audio available in wav or raw.
Limits
Chunking multiplies requests, and each is billed on its characters, so the base total is the same as one big job (the default 5.5% agent fee applies on top of the rate), but a failure costs less to retry. Only paid, completed jobs capture credit. See joining voiceover takes and the API reference for the field list.
Related posts
More in Developers
- TTS language omitted: Spanish text comes out with an English default
On Sume TTS, an omitted language defaults to English at the provider, with a ko or ja fallback only for Hangul or kana text. Set language for all others.
- TTS sentence slices have no audio_url: emit_audio needs wav or raw
With segmentation on, Sume TTS only returns a sample-exact audio_url per sentence for wav or raw output. For mp3 you get timings and a warning instead.
- TTS speed: slow, normal, fast or generation_config.speed 0.6 to 1.5?
Sume TTS marks the slow, normal and fast speed enum deprecated. Send generation_config.speed from 0.6 to 1.5, plus volume 0.5 to 2 and an emotion string.
- TTS call returned processing in sync mode: poll, do not resubmit
Sume TTS in sync mode waits at most 30 seconds. If the job is not terminal, poll status_url, and retry a submit only with the same Idempotency-Key.
Written by Sume