Mercury Voice p95 750 ms: one turn in twenty is slower

Inception reports a 320 ms median and 750 ms p95 for Mercury Voice. At p95, about one turn in 20 is slower. What that means over a 10-turn call.

4 min readSume
All posts

What does a 750 ms p95 mean for a phone call? Inception reports that Mercury Voice returns its first answer token in under 320 milliseconds at the median and 750 milliseconds at p95. A p95 of 750 ms means about 5 percent of turns take longer than that to start, or about one in twenty.

Over a 10-turn call, if turns were independent, the chance that at least one turn exceeds 750 ms is about 40 percent. This is arithmetic on the vendor's figure, not a measurement of the product.

The arithmetic

If each turn has a 5 percent chance of being slower than the p95, then the chance that all n turns are faster is 0.95 to the power n. The chance of at least one slow turn is one minus that. Real turns are not fully independent, since a long prompt or a busy hour affects several in a row, so treat the table as a rough guide.

Chance of at least one turn slower than a p95 threshold, calculated (read 2026-10-03)
Turns in the callChance all turns are under p95Chance of at least one slower turn
195.0%5.0%
577.4%22.6%
1059.9%40.1%
2035.8%64.2%

What the figure covers and does not

The figure is time to the first answer token from the language model. It does not include hearing the caller, deciding that the caller has finished speaking, or turning text into audio. Microsoft's streaming transcriber reports partial text in just over 100 ms, per its announcement, and speech output adds its own start-up time.

So the delay a caller feels is a sum of several stages, each with its own tail. A slow turn is more likely to come from one stage being slow than from all of them at once.

What to do with the tail

Design for the slow turn, not the median.

  • Test with your own prompts and record the median and the 95th percentile separately, from input received to first audio out.
  • Cover the gap with a short acknowledgement phrase on turns that look slow, so silence does not read as a dropped call.
  • Keep a timeout and a fallback response, so a stalled model call does not leave the line dead.
  • Log slow turns with their prompt length, since long system prompts are a common cause.

Where non-live work goes

Not every audio task has a caller waiting. Narration, captions and music for a finished video can run as jobs, with a tail that nobody hears. A Sume job is read by polling or webhook, and its synchronous wait is capped at 30 seconds, per Sume jobs and results. Keep the strict latency work for the live path and let batch audio take the time it needs.

Remember that Mercury Voice is enterprise-only at launch, so the first step for most teams is a conversation with sales, covered in what to ask.

Sources

Related posts

More in Models

All Models posts

Written by Sume