Grok TTS takes 60,000 characters; Sume takes 20,000: how to split
Grok TTS takes 60,000 characters per request; Sume TTS takes 20,000. How to split a long script into jobs and join the audio without gaps.

Which limit is bigger, Grok or Sume?
xAI is three times larger per request: its docs say REST and server-streamed requests take a maximum of 60,000 characters, while a Sume TTS 1.0 transcript is 1 to 20,000 characters, with spaces and punctuation counted. Sume adds a second ceiling that xAI does not mention on that page: synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured.
So a 50,000-character script that fits in one Grok call needs at least three Sume calls, and in practice more, because 1,200 seconds is the limit that usually bites first on slow, careful narration.
What does the xAI page say about WebSocket?
The WebSocket endpoint has no overall limit, but each individual message is capped at 60,000 characters, and the page lists up to 50 concurrent sessions per team. That is a streaming design: you keep a socket open and send text as it arrives.
Sume TTS is deliberately not that. The OpenAPI description calls it an async job plus poll or webhook, non-streaming. A job returns audio artifacts when it finishes, and the sync and subscribe modes wait at most 30 seconds before handing you polling URLs.
| Limit | xAI TTS | Sume TTS 1.0 |
|---|---|---|
| Characters per request | 60,000 (REST) | 20,000 |
| Audio length ceiling | Not stated on the page | 1,200 s, then tts_duration_exceeded |
| Delivery | REST or WebSocket stream | Async job, poll or webhook |
| Longest blocking wait | Not applicable | 30 s (sync / subscribe) |
How do you split a long script for Sume without cutting mid-sentence?
Split on sentence boundaries, keep each part around 12,000 characters or less, and give every part its own stable Idempotency-Key so a retry returns the original job instead of billing twice. The Python below does that with the standard library only.
import json, os, re, urllib.request
def chunks(text, limit=12000):
out, cur = [], ""
for s in re.split(r"(?<=[.!?])\s+", text):
if cur and len(cur) + len(s) + 1 > limit:
out.append(cur)
cur = ""
cur = f"{cur} {s}".strip()
return out + [cur] if cur else out
def submit(part, n, handle):
body = {"transcript": part, "avatar_handle": handle, "mode": "async"}
req = urllib.request.Request(
"https://api.sume.com/v1/tts-1.0/generate",
json.dumps(body).encode(),
{"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Content-Type": "application/json",
"Idempotency-Key": f"audiobook-ch1-part-{n:02d}"})
with urllib.request.urlopen(req) as r:
return json.load(r)
for n, part in enumerate(chunks(open("chapter1.txt").read()), 1):
print(n, len(part), submit(part, n, "narrator"))Why does the 1,200-second ceiling matter more than the character cap?
Speech runs at roughly 15 characters per second for a measured narration pace, which is a rule of thumb from about 150 words a minute at six characters a word, not a Sume measurement. At that rate 1,200 seconds holds about 18,000 characters, so a part filled to the 20,000-character cap can fail on duration even though the character check passes.
Two things change the real number: the pace you set with generation_config.speed (0.6 to 1.5) and the amount of punctuation and numbers in the script. Slow a read to 0.8 and the same text runs about a quarter longer. So pick a part size with headroom, run one part first, read duration_seconds from its result, and size the remaining parts from that.
Because a failed duration check captures no credit, a too-large part costs you time rather than money, but it still wastes a queue slot. Splitting at 12,000 characters, as the script does, leaves room even for a slow pace.
How do you join the parts, and what stays your job?
Each part comes back as its own audio artifact. Sume's timeline audio concat joins up to 20 Sume-hosted parts into one gapless file at the sample level, with no re-synthesis. The timeline audio page notes that its mp3 output re-adds priming padding at every edge, so keep wav when a file will be joined again.
What Sume does not do: it does not split your text for you, and it does not stream audio while you type. If you need live, token-by-token speech for an agent, xAI's WebSocket or another realtime API is the better fit, as the realtime versus async post explains.
Keep the chapter size in mind as well. The next section shows why 20,000 characters is not a safe part size, and the audiobook chapter post covers chapter-sized jobs.
Sources
Related posts
More in Comparisons
- Grok TTS codecs and sample rates vs Sume TTS output_format
xAI TTS and Sume TTS offer the same six sample rates and mp3 bit-rate range. They differ on defaults (24 kHz vs 44.1 kHz) and Sume adds raw and float PCM.
- Grok TTS language "auto" vs Sume's explicit language field
xAI TTS can auto-detect the language of your text. Sume TTS cannot: omit language and Spanish is read as English, except Korean and Japanese.
- Grok TTS speech tags like laugh and whisper vs Sume's emotion field
xAI lets you write [laugh] and wrap text in whisper or singing tags. Sume TTS documents no tag grammar, only emotion, speed and volume in generation_config.
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
Written by Sume