TTS latency numbers side by side: Eleven v4 Turbo, Voxtral, MAI Flash

Three vendors, three latency figures, three different things measured. A table of what each page says, and why a Sume TTS job is a different question.

5 min readSume
All posts

Which text-to-speech model is fastest? The three newest launches quote three numbers that are not the same measurement, so a ranking built from them would be wrong. This table lines up what each vendor's own page says, then explains where a Sume TTS job sits.

Latency matters for a voice agent that must start talking while a person waits. For narrating a video it mostly does not, because nobody is waiting on the first syllable.

What each vendor wrote

Read the middle column carefully: it is the vendor's wording, not a normalized benchmark.

Latency figures as stated on vendor pages, read 2026-10-04
ModelFigureHow the page words it
Eleven v4~100 msMedian inference latency (ElevenLabs)
Eleven v4 Turbo~150 msMedian time to first speech (ElevenLabs)
Voxtral TTS70 msModel latency for typical inputs (Mistral)
MAI-Voice-2.1~550 msModel inference (Microsoft model page)
MAI-Voice-2.1-Flash~45 msModel inference (Microsoft model page)
MAI-Voice-2.1-Flash150 msEnd-to-end, for 45 seconds of audio (Microsoft announcement)

Why the numbers do not rank

Inference latency is how long the model takes once it has the text. Time to first speech adds the path to your client. End-to-end figures depend on how much audio is generated. The same vendor, Microsoft, publishes two different Flash figures for those reasons. Treat each as a claim about its own definition and measure the one you care about on your network.

A fair test sends the same ten short lines to each candidate from the region where your app runs, records the time to the first audio byte and the time to the last, and reports the median and the slowest.

Where a Sume TTS job sits

Sume's TTS 1.0 is an asynchronous job: you submit text, then poll or receive a webhook, and read a hosted audio file. The contract calls phase 1 non-streaming. mode: sync is an alias for a bounded wait of at most 30 seconds, and if the job is still running you poll rather than resubmit.

So the question for Sume is time to a finished file, not time to first sound. That suits narration, voice-overs and anything that goes into a render. It does not suit a live phone agent, which needs a streaming endpoint. The job envelope is described in Jobs and results.

If you are choosing between a live voice model and a job for a narrated video, read realtime model or async TTS job.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume