Eleven v4 Turbo 150 ms first speech vs Sume async TTS jobs

ElevenLabs quotes about 150 ms to first speech for v4 Turbo. Sume TTS is an async job, built for finished narration. Which one fits your use?

5 min readSume
All posts

Eleven v4 Turbo and Sume TTS are built for different jobs. ElevenLabs' page quotes about 150 ms to first speech for v4 Turbo, which suits a live voice agent. Sume TTS 1.0 is an asynchronous job API that returns a finished audio file, with an optional bounded wait of at most 30 seconds. If your product speaks to a caller in real time, latency to first audio is the number that matters. If it produces voiceovers, it is not.

This is a comparison of shape, not a speed race, and Sume makes no latency claim here.

What does ElevenLabs say?

On the Eleven v4 page, read 2026-10-04, ElevenLabs lists v4 Turbo at roughly 150 ms time to first speech and about 100 ms median inference latency. It lists a single generation of up to 10,000 characters, support for more than 90 languages, inline tags such as [pause] and [whispers], and output as MP3, WAV or PCM, and u-law.

Latency and shape, vendor pages read 2026-10-04
AspectEleven v4 (ElevenLabs page)Sume TTS 1.0 (OpenAPI spec)
DeliveryREST, streaming endpoints, SDKsAsync job, polling URLs or webhook
Quoted time to first speechAbout 150 ms for v4 TurboNot quoted
Characters per requestUp to 10,000Up to 20,000
Output containersMP3, WAV or PCM, u-lawmp3, wav, raw

When does the job model win?

When you need a finished file and a record. A Sume job has an id, a status, events and a result you can fetch again, and retries can be made safe with an Idempotency-Key. Audio over 1,200 seconds fails with tts_duration_exceeded, so long scripts are split into several jobs and joined afterwards.

  • Voiceovers for videos and ads, where nobody is waiting on the first syllable.
  • Batches of scripts run overnight.
  • Pipelines that need a webhook when audio is ready.

What does a Sume call look like?

Submit, then poll until the job is terminal:

import os, time, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"

r = requests.post(B + "/v1/tts-1.0/generate", headers=H, timeout=60, json={
    "transcript": "Welcome back. Here is this week's update.",
    "voice": {"id": os.environ["VOICE_ID"]},
    "language": "en",
})
r.raise_for_status()
job_id = r.json()["data"]["job"]["id"]

while not requests.get(B + f"/v1/jobs/{job_id}/status", headers=H,
                       timeout=30).json()["data"]["terminal"]:
    time.sleep(2)
print(job_id)

What should you decide?

Live conversation: use a streaming vendor. Produced audio: use jobs. Compare the limits in Eleven v4's 10,000 characters vs Sume's 20,000 and language coverage in Eleven v4 90 languages vs Sonic 3.6.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume