ElevenLabs previous_text and next_text vs Sume TTS: split scripts

ElevenLabs lets you pass previous_text and next_text for prosody continuity. Sume TTS has no such fields, so here is how to keep long scripts smooth.

5 min readSume
All posts

What the ElevenLabs fields are for

The ElevenLabs text to speech page describes previous_text and next_text as a way to give the model context around the passage you are rendering, so that prosody stays continuous when you generate a long script in separate requests. If you split a chapter into pieces, each piece knows what came before and after.

That is a real feature for chunked workflows, because each separate render otherwise starts and ends as if it were a standalone clip.

What Sume TTS 1.0 does instead

The Sume TTS 1.0 request in the API reference has a transcript field of up to 20,000 characters and no neighbour-context fields. The design answer is to make fewer splits: one request can cover a whole script as long as the audio stays within 1,200 seconds.

Fewer seams means fewer places for the delivery to change. A single 20-minute read has one continuous pass, where five chunks stitched together have four joins.

Chunking options for a long script (read 2026-10-03)
ApproachSeamsProsody continuity
One Sume TTS 1.0 request, up to 20,000 characters and 1,200 s0Whole script in one pass
Several Sume requests split at paragraph breaks1 per splitNo neighbour context; keep splits at natural pauses
ElevenLabs with previous_text and next_text1 per splitContext passed to each request

Splitting well when you must

Some scripts are longer than one request can hold, and then the cut location is the lever you have. Put cuts where a human reader would breathe: between paragraphs, after a heading, or after a full stop that ends a thought.

  • Cut on paragraph boundaries, never mid-sentence.
  • Keep the same voice, language and generation_config on every chunk so only the text differs.
  • Ask for sentence segmentation with word timestamps to learn exactly where each sentence ends and check the seam. The boundary lead is 70 ms by default.
  • Give each chunk its own Idempotency-Key so a retry returns the original job.

Joining the pieces

Join the files losslessly if you can. Ask for wav, concatenate with Timeline audio, and only encode to mp3 at the end. Joining mp3 files adds padding at each edge, which shows up as a tiny gap at every seam.

The cost is the same however you split: Sume prices per character, so five chunks cost the same as one request of the same total text.

A test that shows whether seams matter

Run it yourself. Take a 3,000-character script with a clear change of mood halfway. Render it once as a single request, then again as two requests split at the mood change, with the same voice and settings. Listen at the join, with headphones, at the speed your audience will hear it.

If you cannot hear the seam, split freely and spend your time elsewhere. If you can, keep the script in one request wherever the 1,200-second limit allows.

Takeaway

If you need ElevenLabs-style neighbour context specifically, that is a field Sume does not offer. If you are chasing a smooth long read, send as much as one request allows and join the rest with wav.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume