Context stitching and streaming TTS: Eleven v4 vs Sume async jobs
Eleven v4 lists context stitching and bidirectional streaming. Sume TTS is async and non-streaming, so keep a section in one job and join sections.

Eleven v4 lists context stitching and bidirectional streaming among its features, which suit live or chunked narration. Sume TTS 1.0 is an async job with poll or webhook delivery and no streaming, so the Sume equivalent is to keep each section in one request and join sections afterward.
What does each side say?
Both rows come from each provider's own page or schema, read on 2026-10-09. The Eleven page names the features; it does not describe how they behave, so this post does not either.
| Item | Eleven v4 | Sume TTS 1.0 |
|---|---|---|
| Chunk continuity | Context stitching listed | Not offered; one job is one continuous take |
| Streaming | Bidirectional streaming listed | Async job, non-streaming (OpenAPI: phase 1 is async job plus poll or webhook) |
| Single request size | 10,000 characters per generation | 20,000 characters per request |
| Where documented | elevenlabs.io/v4 | Sume OpenAPI |
How do I keep the voice steady across a long script?
Prosody is decided over the whole text of a request. Splitting a paragraph across two requests asks the voice to start fresh in the middle of an idea. With Sume, put whole scenes in one job (up to 20,000 characters) and split only at section breaks, where a change of pace is natural.
When you need per-sentence control after the fact, request timestamps: { words: true } and segmentation: { mode: "sentence" } with a wav container. You get gapless sentence slices from the same take, so editing a line out does not mean re-synthesizing its neighbours.
When is a streaming API the better choice?
Live agents and low-latency calls need streaming, and Sume's TTS endpoint is the wrong shape for them. For finished narration, ads and course audio the async shape is fine, and a job result is a durable Sume-hosted file you can reuse.
Join sections with the flat $0.01 concat from Timeline audio: up to 20 ordered parts, no silence added at the seams.
Sources
Related posts
More in Comparisons
- Deepgram Flux Multilingual $0.0078/min: 7,500 minutes vs Sume STT $75
Deepgram lists Flux Multilingual streaming at $0.0078 a minute. For 7,500 minutes that is $58.50, against $75.00 on Sume STT. The full Deepgram ladder is below.
- ElevenLabs Scribe v2 $0.22/hour vs Sume STT $0.60/hour: 7.5 hours
Scribe v2 lists $0.22 per hour and Sume STT is $0.01 per minute, or $0.60 per hour. For 7.5 hours of audio: $1.65 vs $4.50, 2.7x apart.
- Scribe v2 Realtime $0.39 vs batch $0.22: 37 hours vs Sume STT
ElevenLabs charges 77 percent more for Scribe v2 Realtime than Scribe v2. For 37 hours the gap is $14.43 against $8.14, and Sume STT is $22.20.
- Scribe v2 for a 6-hour workshop with both add-ons vs Sume STT
Six hours of audio on ElevenLabs Scribe v2 is $1.32 base and $2.04 with keyterm prompting and transcript editing. Sume STT at $0.01 a minute is $3.60.
Written by Sume