Grok TTS takes 60,000 characters; Sume takes 20,000: how to split

Grok TTS takes 60,000 characters per request; Sume TTS takes 20,000. How to split a long script into jobs and join the audio without gaps.

5 min readSume
All posts

Which limit is bigger, Grok or Sume?

xAI is three times larger per request: its docs say REST and server-streamed requests take a maximum of 60,000 characters, while a Sume TTS 1.0 transcript is 1 to 20,000 characters, with spaces and punctuation counted. Sume adds a second ceiling that xAI does not mention on that page: synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured.

So a 50,000-character script that fits in one Grok call needs at least three Sume calls, and in practice more, because 1,200 seconds is the limit that usually bites first on slow, careful narration.

What does the xAI page say about WebSocket?

The WebSocket endpoint has no overall limit, but each individual message is capped at 60,000 characters, and the page lists up to 50 concurrent sessions per team. That is a streaming design: you keep a socket open and send text as it arrives.

Sume TTS is deliberately not that. The OpenAPI description calls it an async job plus poll or webhook, non-streaming. A job returns audio artifacts when it finishes, and the sync and subscribe modes wait at most 30 seconds before handing you polling URLs.

Per-request limits, xAI docs and Sume OpenAPI (read 2026-10-02)
LimitxAI TTSSume TTS 1.0
Characters per request60,000 (REST)20,000
Audio length ceilingNot stated on the page1,200 s, then tts_duration_exceeded
DeliveryREST or WebSocket streamAsync job, poll or webhook
Longest blocking waitNot applicable30 s (sync / subscribe)

How do you split a long script for Sume without cutting mid-sentence?

Split on sentence boundaries, keep each part around 12,000 characters or less, and give every part its own stable Idempotency-Key so a retry returns the original job instead of billing twice. The Python below does that with the standard library only.

import json, os, re, urllib.request

def chunks(text, limit=12000):
    out, cur = [], ""
    for s in re.split(r"(?<=[.!?])\s+", text):
        if cur and len(cur) + len(s) + 1 > limit:
            out.append(cur)
            cur = ""
        cur = f"{cur} {s}".strip()
    return out + [cur] if cur else out

def submit(part, n, handle):
    body = {"transcript": part, "avatar_handle": handle, "mode": "async"}
    req = urllib.request.Request(
        "https://api.sume.com/v1/tts-1.0/generate",
        json.dumps(body).encode(),
        {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
         "Content-Type": "application/json",
         "Idempotency-Key": f"audiobook-ch1-part-{n:02d}"})
    with urllib.request.urlopen(req) as r:
        return json.load(r)

for n, part in enumerate(chunks(open("chapter1.txt").read()), 1):
    print(n, len(part), submit(part, n, "narrator"))

Why does the 1,200-second ceiling matter more than the character cap?

Speech runs at roughly 15 characters per second for a measured narration pace, which is a rule of thumb from about 150 words a minute at six characters a word, not a Sume measurement. At that rate 1,200 seconds holds about 18,000 characters, so a part filled to the 20,000-character cap can fail on duration even though the character check passes.

Two things change the real number: the pace you set with generation_config.speed (0.6 to 1.5) and the amount of punctuation and numbers in the script. Slow a read to 0.8 and the same text runs about a quarter longer. So pick a part size with headroom, run one part first, read duration_seconds from its result, and size the remaining parts from that.

Because a failed duration check captures no credit, a too-large part costs you time rather than money, but it still wastes a queue slot. Splitting at 12,000 characters, as the script does, leaves room even for a slow pace.

How do you join the parts, and what stays your job?

Each part comes back as its own audio artifact. Sume's timeline audio concat joins up to 20 Sume-hosted parts into one gapless file at the sample level, with no re-synthesis. The timeline audio page notes that its mp3 output re-adds priming padding at every edge, so keep wav when a file will be joined again.

What Sume does not do: it does not split your text for you, and it does not stream audio while you type. If you need live, token-by-token speech for an agent, xAI's WebSocket or another realtime API is the better fit, as the realtime versus async post explains.

Keep the chapter size in mind as well. The next section shows why 20,000 characters is not a safe part size, and the audiobook chapter post covers chapter-sized jobs.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume