Eleven v4 Turbo 150 ms first speech vs Sume async TTS jobs
ElevenLabs quotes about 150 ms to first speech for v4 Turbo. Sume TTS is an async job, built for finished narration. Which one fits your use?

Eleven v4 Turbo and Sume TTS are built for different jobs. ElevenLabs' page quotes about 150 ms to first speech for v4 Turbo, which suits a live voice agent. Sume TTS 1.0 is an asynchronous job API that returns a finished audio file, with an optional bounded wait of at most 30 seconds. If your product speaks to a caller in real time, latency to first audio is the number that matters. If it produces voiceovers, it is not.
This is a comparison of shape, not a speed race, and Sume makes no latency claim here.
What does ElevenLabs say?
On the Eleven v4 page, read 2026-10-04, ElevenLabs lists v4 Turbo at roughly 150 ms time to first speech and about 100 ms median inference latency. It lists a single generation of up to 10,000 characters, support for more than 90 languages, inline tags such as [pause] and [whispers], and output as MP3, WAV or PCM, and u-law.
| Aspect | Eleven v4 (ElevenLabs page) | Sume TTS 1.0 (OpenAPI spec) |
|---|---|---|
| Delivery | REST, streaming endpoints, SDKs | Async job, polling URLs or webhook |
| Quoted time to first speech | About 150 ms for v4 Turbo | Not quoted |
| Characters per request | Up to 10,000 | Up to 20,000 |
| Output containers | MP3, WAV or PCM, u-law | mp3, wav, raw |
When does the job model win?
When you need a finished file and a record. A Sume job has an id, a status, events and a result you can fetch again, and retries can be made safe with an Idempotency-Key. Audio over 1,200 seconds fails with tts_duration_exceeded, so long scripts are split into several jobs and joined afterwards.
- Voiceovers for videos and ads, where nobody is waiting on the first syllable.
- Batches of scripts run overnight.
- Pipelines that need a webhook when audio is ready.
What does a Sume call look like?
Submit, then poll until the job is terminal:
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"
r = requests.post(B + "/v1/tts-1.0/generate", headers=H, timeout=60, json={
"transcript": "Welcome back. Here is this week's update.",
"voice": {"id": os.environ["VOICE_ID"]},
"language": "en",
})
r.raise_for_status()
job_id = r.json()["data"]["job"]["id"]
while not requests.get(B + f"/v1/jobs/{job_id}/status", headers=H,
timeout=30).json()["data"]["terminal"]:
time.sleep(2)
print(job_id)
What should you decide?
Live conversation: use a streaming vendor. Produced audio: use jobs. Compare the limits in Eleven v4's 10,000 characters vs Sume's 20,000 and language coverage in Eleven v4 90 languages vs Sonic 3.6.
Sources
Related posts
More in Comparisons
- ElevenLabs Music for a TV spot: two pages that do not agree
ElevenLabs music page says self-serve plans exclude film, TV and games; the API docs say cleared for film and TV. Get this in writing before a broadcast ad.
- ElevenLabs Music terms: restricted industries and banned inputs
Before an ad brief goes into ElevenLabs Music: the terms list restricted industries and prompt inputs you cannot send. What Sume does and does not enforce.
- ElevenLabs TTS on fal at $0.10 per 1,000 characters vs direct and Sume
fal lists ElevenLabs TTS at $0.10 per 1,000 characters, ElevenLabs lists $0.08 for v3, Sume is $0.0475. A 2,000-character script: $0.20, $0.16, $0.095.
- fal image model list vs the Sume catalog: how to check by query
fal.ai lists Seedream 5.0, GPT Image 2.5, Flux 2, Nano Banana 2, Ideogram 4 and Krea 2. How to see which ones a Sume key can call, with a short Python script.
Written by Sume