Split a narration script by model character limit: Python

ElevenLabs lists a 40,000 character limit for Flash v2.5 and 10,000 for v4 and v4 Turbo. A Python splitter that cuts on sentences, with a concat step on Sume.

5 min readSume
All posts

Cut a long script into chunks that stay under the model's character limit, split on sentence ends so no sentence is cut in half, then join the audio afterwards. The ElevenLabs models page, read 2026-10-03, lists a 40,000 character limit for Flash v2.5 (eleven_flash_v2_5) and 10,000 characters for v4 and v4 Turbo. The code below does the splitting.

The limits

Limits come from the vendor models page.

ElevenLabs model limits (read 2026-10-03)
ModelLanguagesCharacter limitLatency note
Flash v2.5 (eleven_flash_v2_5)3240,000about 75 ms
v490+10,000not stated here
v4 Turbo90+10,000not stated here
Multilingual v229not stated herenot stated here

A sentence-safe splitter

This function keeps each chunk under a limit, and raises an error for a sentence that is itself longer than the limit. It uses only the standard library and runs as written.

import re

def split_script(text: str, limit: int) -> list[str]:
    if limit < 1:
        raise ValueError("limit must be positive")
    sentences = re.split(r"(?<=[.!?])\s+", text.strip())
    chunks, current = [], ""
    for s in sentences:
        if len(s) > limit:
            raise ValueError(f"one sentence is {len(s)} chars, over {limit}")
        candidate = f"{current} {s}".strip()
        if len(candidate) <= limit:
            current = candidate
        else:
            chunks.append(current)
            current = s
    if current:
        chunks.append(current)
    return chunks

if __name__ == "__main__":
    script = "One. " * 5
    print([len(c) for c in split_script(script, 12)])

Why chunk below the limit

The limit is a ceiling. In a video, shorter chunks are easier to re-take, because a bad read costs one chunk rather than a whole script. A chunk per scene also lets you place each line against its shot. Leave headroom under the limit rather than packing to the exact count.

Edge cases

Abbreviations such as Dr. or e.g. end with a period and will split a sentence in the wrong place. For scripts with many of them, mark scene breaks yourself with a blank line and split on those first, then apply the sentence splitter inside each scene.

Numbers and URLs are another source of surprises. Spell numbers the way you want them read, since that is what the speech model receives.

Stitching the audio

On Sume, speech is created through the hosted MCP tts_create tool; the tools doc suggests one tts_create per sentence inside script_run when you need many. Once you have several Sume-hosted audio files, timeline audio concatenates 1-20 ordered parts into a single gapless file at $0.01 flat per job. Parts must share a channel layout, and the produced audio is capped at 1800 seconds. Keep wav while the file will be joined again, since mp3 adds priming padding at every edge.

The result includes segments[] with the start offset of each part, which you can use to re-base shot timings.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume