Context stitching and streaming TTS: Eleven v4 vs Sume async jobs

Eleven v4 lists context stitching and bidirectional streaming. Sume TTS is async and non-streaming, so keep a section in one job and join sections.

4 min readSume
All posts

Eleven v4 lists context stitching and bidirectional streaming among its features, which suit live or chunked narration. Sume TTS 1.0 is an async job with poll or webhook delivery and no streaming, so the Sume equivalent is to keep each section in one request and join sections afterward.

What does each side say?

Both rows come from each provider's own page or schema, read on 2026-10-09. The Eleven page names the features; it does not describe how they behave, so this post does not either.

Delivery model (read 2026-10-09)
ItemEleven v4Sume TTS 1.0
Chunk continuityContext stitching listedNot offered; one job is one continuous take
StreamingBidirectional streaming listedAsync job, non-streaming (OpenAPI: phase 1 is async job plus poll or webhook)
Single request size10,000 characters per generation20,000 characters per request
Where documentedelevenlabs.io/v4Sume OpenAPI

How do I keep the voice steady across a long script?

Prosody is decided over the whole text of a request. Splitting a paragraph across two requests asks the voice to start fresh in the middle of an idea. With Sume, put whole scenes in one job (up to 20,000 characters) and split only at section breaks, where a change of pace is natural.

When you need per-sentence control after the fact, request timestamps: { words: true } and segmentation: { mode: "sentence" } with a wav container. You get gapless sentence slices from the same take, so editing a line out does not mean re-synthesizing its neighbours.

When is a streaming API the better choice?

Live agents and low-latency calls need streaming, and Sume's TTS endpoint is the wrong shape for them. For finished narration, ads and course audio the async shape is fine, and a job result is a durable Sume-hosted file you can reuse.

Join sections with the flat $0.01 concat from Timeline audio: up to 20 ordered parts, no silence added at the seams.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume