Split a narration script by model character limit: Python
ElevenLabs lists a 40,000 character limit for Flash v2.5 and 10,000 for v4 and v4 Turbo. A Python splitter that cuts on sentences, with a concat step on Sume.

Cut a long script into chunks that stay under the model's character limit, split on sentence ends so no sentence is cut in half, then join the audio afterwards. The ElevenLabs models page, read 2026-10-03, lists a 40,000 character limit for Flash v2.5 (eleven_flash_v2_5) and 10,000 characters for v4 and v4 Turbo. The code below does the splitting.
The limits
Limits come from the vendor models page.
| Model | Languages | Character limit | Latency note |
|---|---|---|---|
Flash v2.5 (eleven_flash_v2_5) | 32 | 40,000 | about 75 ms |
| v4 | 90+ | 10,000 | not stated here |
| v4 Turbo | 90+ | 10,000 | not stated here |
| Multilingual v2 | 29 | not stated here | not stated here |
A sentence-safe splitter
This function keeps each chunk under a limit, and raises an error for a sentence that is itself longer than the limit. It uses only the standard library and runs as written.
import re
def split_script(text: str, limit: int) -> list[str]:
if limit < 1:
raise ValueError("limit must be positive")
sentences = re.split(r"(?<=[.!?])\s+", text.strip())
chunks, current = [], ""
for s in sentences:
if len(s) > limit:
raise ValueError(f"one sentence is {len(s)} chars, over {limit}")
candidate = f"{current} {s}".strip()
if len(candidate) <= limit:
current = candidate
else:
chunks.append(current)
current = s
if current:
chunks.append(current)
return chunks
if __name__ == "__main__":
script = "One. " * 5
print([len(c) for c in split_script(script, 12)])Why chunk below the limit
The limit is a ceiling. In a video, shorter chunks are easier to re-take, because a bad read costs one chunk rather than a whole script. A chunk per scene also lets you place each line against its shot. Leave headroom under the limit rather than packing to the exact count.
Edge cases
Abbreviations such as Dr. or e.g. end with a period and will split a sentence in the wrong place. For scripts with many of them, mark scene breaks yourself with a blank line and split on those first, then apply the sentence splitter inside each scene.
Numbers and URLs are another source of surprises. Spell numbers the way you want them read, since that is what the speech model receives.
Stitching the audio
On Sume, speech is created through the hosted MCP tts_create tool; the tools doc suggests one tts_create per sentence inside script_run when you need many. Once you have several Sume-hosted audio files, timeline audio concatenates 1-20 ordered parts into a single gapless file at $0.01 flat per job. Parts must share a channel layout, and the produced audio is capped at 1800 seconds. Keep wav while the file will be joined again, since mp3 adds priming padding at every edge.
The result includes segments[] with the start offset of each part, which you can use to re-base shot timings.
Sources
Related posts
More in Developers
- Patching Supabase Postgres 17.11 vs Sume's 10-attempt webhook budget
Supabase's September 25 Postgres 15.19 and 17.11 releases fix 44 CVEs. A restart can outlast Sume's ten 30-second webhook attempts, so plan a redeliver.
- Supabase cached egress is $0.03/GB: cost of serving a 20 MB AI clip
Supabase lists cached Storage egress at $0.03 per GB. Worked arithmetic for serving generated clips, and when to link a Sume media URL instead of copying.
- End-user id on jobs: OpenAI safety identifier vs Sume metadata
OpenAI's Realtime guide asks for an OpenAI-Safety-Identifier header. Sume stores caller metadata on the job but does not send it to the provider. Use both.
- Test a faster-and-cheaper claim with your own timings and usage.cost
Luma's news page says Ray3.14 is 4x faster and 3x cheaper. A short Python script turns your own job timings and usage.cost values into two ratios you can trust.
Written by Sume