Flash TTS '55% faster inference' vs the end-to-end time of a TTS job
Microsoft quotes 55% faster inference and 150 ms end to end for MAI-Voice-2.1-Flash. Sume TTS is an async job; here is what you can actually time.

Microsoft describes MAI-Voice-2.1-Flash with two speed figures: 55% faster inference and 150 ms end to end. They measure different things, and neither is the time a Sume TTS job takes. Sume's TTS is an asynchronous job, not a stream: you submit text, then wait for a terminal state and fetch a finished file. The honest comparison for a voiceover is the full job time you measure on your own scripts.
Three clocks, not one
Inference time is the model's compute. End-to-end time in a streaming API is usually the delay to first audio. Job time on Sume is submit to finished result, including queueing, generation and storing the file. A faster inference number can lower job time, but queueing and result handling do not shrink with it. Sume's sync mode waits for at most 30 seconds (wait_timeout_seconds 0-30) and returns the envelope with a sync.timed_out flag, so a short line often comes back in one call, and a long one needs polling.
| Figure | Source | Measures |
|---|---|---|
| 55% faster inference | MAI-Voice-2.1-Flash announcement | Model compute, versus a prior baseline |
| 150 ms end to end | MAI-Voice-2.1-Flash announcement | Microsoft quotes it for the request path |
| Job time | Sume TTS router | Submit to terminal state, measured by you |
| Sync wait | Sume docs | Up to 30 s, then poll; do not resubmit |
Time your own jobs
Log the clock around the three events you can see: the submit response, the first terminal status and the result fetch. Run the same 300-character line ten times and take the median, not the minimum. If you generate ahead of playback, as in a prepared video, job time rarely matters. If a person is waiting on a reply spoken aloud, a streaming API is the right shape, and Sume's router lists streaming TTS as a non-goal.
import os, time, requests
H = {"x-api-key": os.environ["SUME_API_KEY"]}
body = {"model": "sonic-3.6", "transcript": "Your table is ready.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"mode": "sync", "wait_timeout_seconds": 30}
t0 = time.time()
d = requests.post("https://api.sume.com/v1/tts-router/generate",
json=body, headers=H, timeout=60).json()["data"]
while not d.get("terminal"):
time.sleep(1)
d = requests.get(d["status_url"], headers=H, timeout=60).json()["data"]
print(round(time.time() - t0, 2), "seconds to terminal")Reading the result
Report a median and a worst case from at least ten runs, and keep the script length with them. A number without a character count cannot be compared with anyone else's figure.
Which workloads care
A prepared video, a podcast read or a batch of product explainers does not care whether first audio arrives in 150 ms. It cares about cost per finished minute and about getting exactly the approved text spoken. A live phone agent cares about first audio and nothing else. Decide which workload you have before you read any speed claim.
For prepared audio, Sume gives you things a latency figure cannot: a stored result URL, a terminal webhook, an Idempotency-Key that stops a retry from billing twice, and, for source-bound jobs, a receipt of the text that was spoken.
A fair way to report your own timing
State the character count, the model id and the mode you used. Report the median of ten runs and the slowest of the ten. Separate the time to submit from the time to terminal, because a slow network affects the first and the provider affects the second. That is enough for someone else to repeat your measurement and compare.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 at $22 per million characters vs Eleven v4
MAI-Voice-2.1 is $22 per 1M characters and Flash $15; Eleven v4 is $22 per 1M in a promo through Oct 12. Sume's Sonic route is $47.50 per 1M characters.
- MAI-Voice styles via express-as vs Sume's emotion field
MAI voices set styles with SSML mstts:express-as; some only have neutral. Sume has no SSML; generation_config.emotion is a free string up to 64 characters.
- Max video length in October 2026: TikTok, YouTube Shorts, Sume tools
TikTok's API allows 10 minutes and YouTube Shorts 3. Sume caps timeline and inspect at 1800 s, trim output at 900 s, and frames and filter at 300 s.
- Meta signs the EU AI labeling code: an advertiser checklist
Meta confirmed on July 28, 2026 it will sign the EU AI-content transparency code. No ad deadline is stated, and AI info labels already exist on ads.
Written by Sume