tts_duration_exceeded: split long scripts for Sume TTS

Sume TTS fails with tts_duration_exceeded when audio would pass 1,200 seconds, and takes at most 20,000 characters. Split rules and how to join the parts.

4 min readSume
All posts

Sume TTS stops a job with tts_duration_exceeded when the synthesized audio would run past 1,200 seconds, which is 20 minutes, and no credit is captured for the failed job. A request also cannot carry more than 20,000 characters in transcript. Split long scripts at paragraph boundaries, generate each part, and join the takes.

The two caps

Sume TTS caps, from the Sume OpenAPI reference, read 2026-10-01. The pace line is an assumption.
CapValueRough equivalent
transcript length20,000 charactersabout 3,300 words
audio length1,200 secondsabout 3,000 words at 150 words per minute
billing$0.0475 per 1,000 characters, plus a 5.5% agent fee by default20,000 characters = $0.95 before the fee

Why you may hit duration first

A slow voice or generation_config.speed of 0.6 stretches speech, so a script well under 20,000 characters can still pass 1,200 seconds. Speed accepts 0.6 to 1.5. If you slow the delivery, shrink the chunk size too.

A splitting recipe

  • Split on blank lines, never mid-sentence.
  • Pack paragraphs until a chunk reaches about 6,000 characters, a comfortable margin below both caps.
  • Use the same model, voice and generation_config for every chunk, or the seams will be audible.
  • Number the files, submit chunks in parallel as async jobs, and wait for all of them with one jobs_wait call (or poll each job's status).
  • Join with Timeline audio concat (POST /v1/timeline-1.0/audio, operation: concat), which joins in the sample domain with no added silence at the seams. It takes 1 to 20 parts per job, so a very long script needs a second pass.

Seam quality

Each chunk starts with fresh intonation, so end chunks at a natural pause, like a heading or a full stop at the end of a paragraph. Listen to each joint once. segmentation on TTS with timestamps.words can give you sentence timing if you need captions, with the slice audio available in wav or raw.

Limits

Chunking multiplies requests, and each is billed on its characters, so the base total is the same as one big job (the default 5.5% agent fee applies on top of the rate), but a failure costs less to retry. Only paid, completed jobs capture credit. See joining voiceover takes and the API reference for the field list.

Related posts

More in Developers

All Developers posts

Written by Sume