ElevenLabs dialogue continuity vs Sume TTS segments: what differs
ElevenLabs Text to Dialogue links takes with previous_text and request ids. Sume TTS has no such fields; it returns gapless sentence segments instead.

Sume TTS has no previous_text, future_text or request-id continuity fields. ElevenLabs added continuity controls to Text to Dialogue, per its changelog (Sep 28) and API reference, read 2026-10-01. Sume instead lets you ask for sentence segments with timing and keep a whole script in one call.
What did ElevenLabs add?
Its reference says previous_text and future_text are each up to 100 characters, and previous_request_ids and next_request_ids take up to 3 ids. These let separately generated takes sound continuous.
| Item | ElevenLabs Text to Dialogue | Sume TTS |
|---|---|---|
| Preceding and following text | previous_text and future_text, up to 100 characters each | No such fields |
| Linking takes by request id | previous_request_ids and next_request_ids, up to 3 ids | No such fields |
| Transcript length | Not covered here | Up to 20,000 characters per call |
| Timed output | Not covered here | Sentence segments with timestamps.segmentation set to sentence |
What does Sume offer?
POST /v1/tts-1.0/generate takes a transcript up to 20,000 characters, so most scripts fit one call and one take. With timestamps.segmentation set to sentence you get gapless sentence segments, and word timestamps are available too. Voice comes from avatar_id, avatar_handle or voice.id.
How do I keep tone across calls?
Keep the same voice and generation_config (speed 0.6 to 1.5, volume 0.5 to 2, an emotion string up to 64 characters) on every call. Sume's docs do not promise that two calls will join without a seam, so listen at the joins.
Which should I pick?
If you need multi-speaker dialogue stitched from many calls, ElevenLabs documents the tools. For one narrator and a timed script, one Sume call is simpler.
Sources
Related posts
More in Models
- Adobe Firefly Image 5 API: native 4 MP vs Sume resolution tiers
Firefly's Image5 model is described as native 4 MP with Instruct Edit. Sume does not list a Firefly model; it uses 512, 1K, 2K and 4K tiers.
- Fish Audio Drama 3 single-word fix vs Sume sentence segments
Fish Audio says Drama 3 preview can fix a single word. Sume has no word-level repair: regenerate a sentence and join takes with Timeline audio concat.
- FLUX 1.1 Ultra raw mode and 4MP: what Sume image size accepts
Raw mode and 4MP belong to BFL's FLUX1.1 [pro] Ultra endpoint. Sume lists FLUX.2 ids and sizes images with aspect_ratio, resolution tiers and image_size.
- FLUX.1 Kontext pro vs FLUX.2 for edits: what to use on Sume
BFL calls FLUX.1 Kontext [pro] previous-generation and recommends FLUX.2 for edits with up to 10 references. On Sume, use flux.2-pro with input_references.
Written by Sume