Text to speech long text limit: 1,200 seconds, not just characters
Sume TTS 1.0 has two limits: 20,000 characters in, 1,200 seconds of audio out. Over the audio limit it fails with tts_duration_exceeded and captures no credit.

The long-text limit for Sume TTS 1.0 is two limits, not one: a transcript of at most 20,000 characters, and synthesized audio of at most 1,200 seconds. Audio longer than 1,200 seconds fails with tts_duration_exceeded, and no credit is captured for the failed job. So chapter an audiobook by how long each piece will speak, not only by how long it is on the page.
Both limits are from the TTS 1.0 request schema in the OpenAPI document behind the API reference, read 2026-09-29. Splitting and joining a whole manuscript is walked through in text to speech for long text and audiobooks; this post is about which limit you hit first.
What are the two limits exactly?
The transcript field takes 1 to 20,000 characters, and the schema notes that spaces and punctuation count toward usage. The output limit is separate and applies to the audio the job produces.
| Limit | Value | What happens past it |
|---|---|---|
| Transcript length | 20,000 characters | Request is invalid |
| Synthesized audio | 1,200 seconds | Fails with tts_duration_exceeded, no credit capture |
Which limit will I hit first?
That depends on how fast the voice speaks, and you can move it: generation_config.speed accepts 0.6 to 1.5, so a slower setting stretches the same text into more seconds. A chapter that fits under 1,200 seconds at speed 1 may not at 0.6. Sume's schema gives no characters-per-second figure, so measure with a sample: synthesize one representative page, read its duration, and size chapters from that with a margin.
How do I cut chapters safely?
Cut at sentence or paragraph ends, keep each chapter well under both limits, and send one request per chapter. Send timestamps.words and segmentation.mode: "sentence" if you want per-sentence timings for each chapter; the segments are gapless.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Chapter one. It was a quiet morning...",
"avatar_handle": "@your-avatar",
"output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
"timestamps": { "words": true }
}'Am I charged when a chapter goes over 1,200 seconds?
Per the schema, no credit is captured when the job fails with tts_duration_exceeded. That makes the failure cheap, but it still costs you a round trip and a re-split, so a duration estimate before submitting is worth having on long chapters. A simple habit is to synthesize the longest chapter of a book first: if it passes, the shorter ones will too at the same speed and voice, and you find out about a problem before you have spent anything on the rest.
How do I join the chapters into one file?
Keep the chapters as wav and join them with Timeline audio, whose concat takes 1 to 20 ordered parts of Sume-hosted audio and joins them in the sample domain with no silence at the seams (see the Timeline audio page).
Sources
Related posts
More in Developers
- Text to speech sentence timestamps API: one lip-sync clip each
Sume TTS 1.0 returns gapless sentence segments; with wav output each gets its own audio_url, ready to drive a lip-sync clip. Request shape and the 5 s catch.
- Let the API pick the video model: sume/auto for vertical UGC clips
Send model sume/auto to POST /v1/videos and Sume picks the family. The response echoes sume/auto and never names the model. When to pin a model instead.
- 400 unknown_parameter on a Format run: read the suggestion
Sume rejects a top-level field it does not know with 400 unknown_parameter and names the likely one. Why webook_url and thread_id fail, and the fix.
- Veo 3.1 personGeneration: allow_adult vs allow_all by mode
Veo 3.1 personGeneration is allow_all for text-to-video, allow_adult for image modes, and allow_adult only in some regions. Sume has no such field.
Written by Sume