Flash TTS '55% faster inference' vs the end-to-end time of a TTS job

Microsoft quotes 55% faster inference and 150 ms end to end for MAI-Voice-2.1-Flash. Sume TTS is an async job; here is what you can actually time.

5 min readSume
All posts

Microsoft describes MAI-Voice-2.1-Flash with two speed figures: 55% faster inference and 150 ms end to end. They measure different things, and neither is the time a Sume TTS job takes. Sume's TTS is an asynchronous job, not a stream: you submit text, then wait for a terminal state and fetch a finished file. The honest comparison for a voiceover is the full job time you measure on your own scripts.

Three clocks, not one

Inference time is the model's compute. End-to-end time in a streaming API is usually the delay to first audio. Job time on Sume is submit to finished result, including queueing, generation and storing the file. A faster inference number can lower job time, but queueing and result handling do not shrink with it. Sume's sync mode waits for at most 30 seconds (wait_timeout_seconds 0-30) and returns the envelope with a sync.timed_out flag, so a short line often comes back in one call, and a long one needs polling.

What each speed figure measures, Microsoft announcement and Sume docs, read 2026-10-05
FigureSourceMeasures
55% faster inferenceMAI-Voice-2.1-Flash announcementModel compute, versus a prior baseline
150 ms end to endMAI-Voice-2.1-Flash announcementMicrosoft quotes it for the request path
Job timeSume TTS routerSubmit to terminal state, measured by you
Sync waitSume docsUp to 30 s, then poll; do not resubmit

Time your own jobs

Log the clock around the three events you can see: the submit response, the first terminal status and the result fetch. Run the same 300-character line ten times and take the median, not the minimum. If you generate ahead of playback, as in a prepared video, job time rarely matters. If a person is waiting on a reply spoken aloud, a streaming API is the right shape, and Sume's router lists streaming TTS as a non-goal.

import os, time, requests
H = {"x-api-key": os.environ["SUME_API_KEY"]}
body = {"model": "sonic-3.6", "transcript": "Your table is ready.",
        "voice": {"id": os.environ["SUME_VOICE_ID"]},
        "mode": "sync", "wait_timeout_seconds": 30}
t0 = time.time()
d = requests.post("https://api.sume.com/v1/tts-router/generate",
                  json=body, headers=H, timeout=60).json()["data"]
while not d.get("terminal"):
    time.sleep(1)
    d = requests.get(d["status_url"], headers=H, timeout=60).json()["data"]
print(round(time.time() - t0, 2), "seconds to terminal")

Reading the result

Report a median and a worst case from at least ten runs, and keep the script length with them. A number without a character count cannot be compared with anyone else's figure.

Which workloads care

A prepared video, a podcast read or a batch of product explainers does not care whether first audio arrives in 150 ms. It cares about cost per finished minute and about getting exactly the approved text spoken. A live phone agent cares about first audio and nothing else. Decide which workload you have before you read any speed claim.

For prepared audio, Sume gives you things a latency figure cannot: a stored result URL, a terminal webhook, an Idempotency-Key that stops a retry from billing twice, and, for source-bound jobs, a receipt of the text that was spoken.

A fair way to report your own timing

State the character count, the model id and the mode you used. Report the median of ten runs and the slowest of the ten. Separate the time to submit from the time to terminal, because a slow network affects the first and the provider affects the second. That is enough for someone else to repeat your measurement and compare.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume