Text to speech long text limit: 1,200 seconds, not just characters

Sume TTS 1.0 has two limits: 20,000 characters in, 1,200 seconds of audio out. Over the audio limit it fails with tts_duration_exceeded and captures no credit.

4 min readSume
All posts

The long-text limit for Sume TTS 1.0 is two limits, not one: a transcript of at most 20,000 characters, and synthesized audio of at most 1,200 seconds. Audio longer than 1,200 seconds fails with tts_duration_exceeded, and no credit is captured for the failed job. So chapter an audiobook by how long each piece will speak, not only by how long it is on the page.

Both limits are from the TTS 1.0 request schema in the OpenAPI document behind the API reference, read 2026-09-29. Splitting and joining a whole manuscript is walked through in text to speech for long text and audiobooks; this post is about which limit you hit first.

What are the two limits exactly?

The transcript field takes 1 to 20,000 characters, and the schema notes that spaces and punctuation count toward usage. The output limit is separate and applies to the audio the job produces.

TTS 1.0 limits for one request, read 2026-09-29.
LimitValueWhat happens past it
Transcript length20,000 charactersRequest is invalid
Synthesized audio1,200 secondsFails with tts_duration_exceeded, no credit capture

Which limit will I hit first?

That depends on how fast the voice speaks, and you can move it: generation_config.speed accepts 0.6 to 1.5, so a slower setting stretches the same text into more seconds. A chapter that fits under 1,200 seconds at speed 1 may not at 0.6. Sume's schema gives no characters-per-second figure, so measure with a sample: synthesize one representative page, read its duration, and size chapters from that with a margin.

How do I cut chapters safely?

Cut at sentence or paragraph ends, keep each chapter well under both limits, and send one request per chapter. Send timestamps.words and segmentation.mode: "sentence" if you want per-sentence timings for each chapter; the segments are gapless.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "transcript": "Chapter one. It was a quiet morning...",
    "avatar_handle": "@your-avatar",
    "output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
    "timestamps": { "words": true }
  }'

Am I charged when a chapter goes over 1,200 seconds?

Per the schema, no credit is captured when the job fails with tts_duration_exceeded. That makes the failure cheap, but it still costs you a round trip and a re-split, so a duration estimate before submitting is worth having on long chapters. A simple habit is to synthesize the longest chapter of a book first: if it passes, the shorter ones will too at the same speed and voice, and you find out about a problem before you have spent anything on the rest.

How do I join the chapters into one file?

Keep the chapters as wav and join them with Timeline audio, whose concat takes 1 to 20 ordered parts of Sume-hosted audio and joins them in the sample domain with no silence at the seams (see the Timeline audio page).

Sources

Related posts

More in Developers

All Developers posts

Written by Sume