TTS latency numbers side by side: Eleven v4 Turbo, Voxtral, MAI Flash
Three vendors, three latency figures, three different things measured. A table of what each page says, and why a Sume TTS job is a different question.

Which text-to-speech model is fastest? The three newest launches quote three numbers that are not the same measurement, so a ranking built from them would be wrong. This table lines up what each vendor's own page says, then explains where a Sume TTS job sits.
Latency matters for a voice agent that must start talking while a person waits. For narrating a video it mostly does not, because nobody is waiting on the first syllable.
What each vendor wrote
Read the middle column carefully: it is the vendor's wording, not a normalized benchmark.
| Model | Figure | How the page words it |
|---|---|---|
| Eleven v4 | ~100 ms | Median inference latency (ElevenLabs) |
| Eleven v4 Turbo | ~150 ms | Median time to first speech (ElevenLabs) |
| Voxtral TTS | 70 ms | Model latency for typical inputs (Mistral) |
| MAI-Voice-2.1 | ~550 ms | Model inference (Microsoft model page) |
| MAI-Voice-2.1-Flash | ~45 ms | Model inference (Microsoft model page) |
| MAI-Voice-2.1-Flash | 150 ms | End-to-end, for 45 seconds of audio (Microsoft announcement) |
Why the numbers do not rank
Inference latency is how long the model takes once it has the text. Time to first speech adds the path to your client. End-to-end figures depend on how much audio is generated. The same vendor, Microsoft, publishes two different Flash figures for those reasons. Treat each as a claim about its own definition and measure the one you care about on your network.
A fair test sends the same ten short lines to each candidate from the region where your app runs, records the time to the first audio byte and the time to the last, and reports the median and the slowest.
Where a Sume TTS job sits
Sume's TTS 1.0 is an asynchronous job: you submit text, then poll or receive a webhook, and read a hosted audio file. The contract calls phase 1 non-streaming. mode: sync is an alias for a bounded wait of at most 30 seconds, and if the job is still running you poll rather than resubmit.
So the question for Sume is time to a finished file, not time to first sound. That suits narration, voice-overs and anything that goes into a render. It does not suit a live phone agent, which needs a streaming endpoint. The job envelope is described in Jobs and results.
If you are choosing between a live voice model and a job for a narrated video, read realtime model or async TTS job.
Sources
Related posts
More in Comparisons
- TTS leaderboard: 33 Elo points rank 5 to 12
On Versely's September 2026 voice leaderboard, ranks 5 to 12 span 33 Elo points. Here is what that gap means and how to test a voice with Sume's tts_create.
- Udio downloads are off: export-ready music for video work
Udio disabled downloads after its UMG settlement. Where to get an exportable AI music bed for video instead: Sume's Music Router returns a file URL.
- Reading a vendor-run TTS leaderboard: Gemini 3.8 Flash TTS on VoiceEQ
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. What that does and does not tell you, plus a blind test to run on your script.
- Veo 3.1 Lite at half of Fast vs an Omni Flash draft
Google says Veo 3.1 Lite costs under half of Veo 3.1 Fast. On Sume the Google video id is gemini-omni-flash-1.1, 3-10 s at 360p to 4K. Compare before budgeting.
Written by Sume