Per-request character limits: Eleven v4 10,000 vs Sume TTS 20,000
Eleven v4 takes up to 10,000 characters per generation, Sume TTS 1 to 20,000. A 45,000-character script is 5 requests versus 3, with a chunking script.

Eleven v4 accepts a single generation of up to 10,000 characters, and the Sume TTS surfaces accept a transcript of 1 to 20,000 characters. For a 45,000-character script that means at least 5 requests on Eleven and 3 on Sume (ceil of 45,000 / 10,000 and of 45,000 / 20,000). Fewer requests means fewer seams to hide, but the limit is only one part of keeping a long read consistent.
The limits side by side
| Item | Eleven v4 | Sume TTS 1.0 / TTS Router |
|---|---|---|
| Text per request | Up to 10,000 characters per generation | 1 to 20,000 characters |
| Text input | Text with inline audio tags | Exactly one of transcript or transcript_source |
| Keeping a long read smooth | Context stitching (product page) | Split by sentence IDs, one job per chunk |
| Proof of submitted text | Not stated on the pages read | transcript_receipt on source-bound jobs: job id, revision, sentence IDs and a SHA-256 of the submitted text |
| Requests for 45,000 characters | 5 | 3 |
Splitting a script on sentence boundaries
Never cut mid-sentence: the seam lands in the middle of a phrase and the intonation resets. This script splits on sentence ends and keeps each chunk under a limit you choose. It splits a single oversized sentence hard, as a last resort.
import asyncio
import re
def chunk(text, limit):
out, cur = [], ''
for s in re.split(r'(?<=[.!?])\s+', text.strip()):
while len(s) > limit:
if cur:
out.append(cur)
cur = ''
out.append(s[:limit])
s = s[limit:]
if cur and len(cur) + len(s) + 1 > limit:
out.append(cur)
cur = s
else:
cur = (cur + ' ' + s).strip()
return out + ([cur] if cur else [])
async def main():
script = 'This is one sentence. ' * 2250 # about 49,500 characters
for limit in (10000, 20000):
parts = chunk(script, limit)
print(limit, len(parts), max(len(x) for x in parts))
asyncio.run(main())Source binding in Sume
For scripted work Sume lets you send transcript_source instead of raw text: a script_revision_id plus ordered sentence_ids from GET /v1/tts-1.0/source. The API resolves the IDs against the accepted revision and rejects duplicates, reordered or unknown IDs. A source-bound job returns a receipt, and the reference states plainly that a receipt proves the submitted text, not the pronunciation. You still listen to the result. Note that GET /v1/tts-1.0/source works on an authenticated thread; for plain scripts you send transcript and chunk it yourself.
Whichever engine you use, render the seams. Play the last two seconds of one chunk into the first two of the next and check pace and level before you approve a long read.
Sources
Related posts
More in Developers
- ElevenLabs Music can sign MP3s with C2PA; what to log for a Sume track
The ElevenLabs Music API has sign_with_c2pa for MP3 output only. Sume's Music docs name no audio credential, so keep a record of job, prompt and routed model.
- Edit a transcript with plain-language instructions: ElevenLabs STT
ElevenLabs STT edits a transcript from an instruction of up to 2,000 characters and returns edited_transcript. Sume captions align your own script_text.
- Embargoed Black Friday reveal: Sume media URLs are public, copy first
Format artifacts sit on durable public media.sume.com URLs with no expiry. For an embargoed drop, copy the file to your own storage and never log the URL.
- Emotion tags in the script or an emotion field: Sume's bill
MAI-Voice takes emotion tags. Sume's documented control is generation_config.emotion, and every transcript character is billed, tags included.
Written by Sume