Eleven v4 Turbo in ElevenAgents vs Sume TTS as async jobs
Eleven v4 Turbo targets live agents. Sume TTS is an async job with poll or webhook and no streaming. Which one fits a call bot and which fits produced audio.

Use Eleven v4 Turbo when a caller is waiting on a spoken reply, and Sume TTS when you are producing audio that a video, ad or app will play later. ElevenLabs says Turbo is live in ElevenAgents, ElevenCreative and the ElevenAPI (read 2026-10-07). Sume's TTS is an asynchronous job you poll or receive by webhook, and its streaming is a stated non-goal.
What ElevenLabs positions Turbo for
The v4 announcement lists a low-latency variant, Eleven v4 Turbo, alongside Eleven v4, with more than 90 languages and voice cloning from about 10 seconds of audio. The page quotes a median time to first speech of about 150 ms for Turbo (read 2026-10-07).
| Question | Eleven v4 Turbo | Sume TTS 1.0 |
|---|---|---|
| Delivery | Low-latency, live agent use | Async job, poll or webhook |
| Streaming | Yes, per the product positioning | Not offered (phase 1 non-streaming) |
| Result | Audio as it is produced | Sume-hosted audio artifact |
| Wait on submit | Not applicable | Optional sync wait up to 30 seconds |
How a Sume TTS job behaves
mode: async returns at once with status_url and result_url. mode: sync and subscribe are the same bounded wait of up to 30 seconds; if the job is still running you get a timed-out flag and keep polling rather than resubmitting. mode: webhook delivers a signed terminal callback, with no partial or progress callbacks.
Synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded, and no credit is captured for that failure. The TTS Router docs also list streaming TTS and an Eleven or OpenAI engine as out of scope for its first ship.
Pick by who is waiting
The third case is the one that fits both worlds. Many agent replies are fixed phrases such as greetings and confirmations, and a pre-rendered file has zero synthesis delay at call time.
- A phone or chat agent that must answer mid-conversation: a streaming engine such as Eleven v4 Turbo.
- A voice line used in a rendered video, an ad or a lesson: a Sume job, with the file saved once and reused.
- A stock of canned replies for a live agent: render them ahead of time as Sume jobs, then play the files.
A pre-render submit
The pcm_mulaw at 8000 Hz output is one of the encodings and rates in the Sume schema, useful for telephony playback.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: greeting-line-001" \
-d '{
"transcript": "Thanks for calling. How can I help?",
"avatar_handle": "speaker",
"language": "en",
"output_format": {"container": "wav", "encoding": "pcm_mulaw", "sample_rate": 8000},
"mode": "webhook",
"webhook_url": "https://example.com/webhooks/sume"
}'Sources
Related posts
More in Comparisons
- ElevenLabs plans vs Sume TTS: cost per million characters
ElevenLabs plans work out to about $0.022 per 1,000 v4 characters, so 1 million characters is about $22. Sume TTS is $47.50 per million.
- FLUX.2 pro or flex for label text on a bottle: Sume ids
Black Forest Labs calls flex the typography and detail model and pro the affordable production model. How to compare them on Sume with one label prompt.
- FLUX 3 Image is on BFL's pricing page: which FLUX ids can Sume call?
bfl.ai/pricing lists FLUX 3 Image beside FLUX.2 pro, flex, max and klein. On Sume you can call FLUX.2 Pro and Flex today. A dated table and the migration rule.
- Gemini 3.5 Transcribe lists 85+ languages; Sume STT takes a hint
Google's changelog lists Gemini 3.5 Transcribe with 85+ languages and word timestamps. Sume STT returns words always and takes an optional language hint.
Written by Sume