Low latency text to speech API: 150 ms streaming vs async jobs

ElevenLabs quotes about 150 ms to first speech for Eleven v4 Turbo. Sume's text to speech returns a finished file per job, so pre-render lines you can predict.

4 min readSume
All posts

A low latency text to speech API starts playing audio within a fraction of a second, which needs streaming. ElevenLabs says Eleven v4 Turbo has a median time to first speech of about 150 ms. Sume's text to speech does not stream: it runs an async job and returns a finished audio file, so it suits lines you can render before the listener needs them, not a live conversation.

The Eleven figure is from ElevenLabs' launch post, read 2026-09-29, and is its own median, not our measurement. The Sume facts are from the TTS 1.0 route in the Sume API reference and Jobs and results.

What is the difference in how audio arrives?

From ElevenLabs' launch post and the Sume TTS 1.0 route, read 2026-09-29.
Eleven v4 Turbo (ElevenLabs' claim)Sume TTS 1.0
First audioMedian about 150 ms to first speechAfter the job finishes
DeliveryNot described in the postAsync job plus poll or webhook, non-streaming
Blocking optionNot described in the postsync waits up to 30 seconds

How long can I make a request wait?

mode: async returns immediately with a status URL and a result URL. mode: sync and mode: subscribe are aliases for one bounded wait of up to wait_timeout_seconds, at most 30. If the job is not done by then, keep polling the same job; do not resubmit.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: greeting-001" \
  -d '{
    "transcript": "Thanks for calling. How can I help?",
    "voice": { "id": "voi_..." },
    "language": "en",
    "mode": "sync",
    "wait_timeout_seconds": 10
  }'

Which lines can I pre-render?

Anything the script knows in advance: greetings, menu prompts, confirmations, narration for a video. Render them ahead, store the files, and play them on demand. Lines that depend on what a person just said cannot be rendered in advance, and for those you need a service that streams.

How do I split a long script?

TTS 1.0 can return word timings and sentence segmentation when you ask for them, and you can send a long script as several short jobs so the first one finishes sooner. Each job is a separate file, so join them with Timeline audio if you need one track. Measure the wait for your own text; the Sume docs give no latency figure.

What does webhook mode add?

Webhook mode delivers exactly three events, job.completed, job.failed and job.canceled, to a public HTTPS webhook_url. There are no progress or partial-audio callbacks, so it does not lower the time to first audio. It helps when a pipeline should continue on its own once a batch of lines is finished, and the docs suggest verifying the signature and polling as a backup.

How do I decide which one I need?

Ask whether the listener is waiting on words you have not seen yet. If a person is talking to the system live and the reply depends on what they said, you need streaming, and ElevenLabs' figure describes that use. If the script is written before playback, an async job that finishes in the background is enough.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume