Mercury Voice p95 750 ms: one turn in twenty is slower
Inception reports a 320 ms median and 750 ms p95 for Mercury Voice. At p95, about one turn in 20 is slower. What that means over a 10-turn call.

What does a 750 ms p95 mean for a phone call? Inception reports that Mercury Voice returns its first answer token in under 320 milliseconds at the median and 750 milliseconds at p95. A p95 of 750 ms means about 5 percent of turns take longer than that to start, or about one in twenty.
Over a 10-turn call, if turns were independent, the chance that at least one turn exceeds 750 ms is about 40 percent. This is arithmetic on the vendor's figure, not a measurement of the product.
The arithmetic
If each turn has a 5 percent chance of being slower than the p95, then the chance that all n turns are faster is 0.95 to the power n. The chance of at least one slow turn is one minus that. Real turns are not fully independent, since a long prompt or a busy hour affects several in a row, so treat the table as a rough guide.
| Turns in the call | Chance all turns are under p95 | Chance of at least one slower turn |
|---|---|---|
| 1 | 95.0% | 5.0% |
| 5 | 77.4% | 22.6% |
| 10 | 59.9% | 40.1% |
| 20 | 35.8% | 64.2% |
What the figure covers and does not
The figure is time to the first answer token from the language model. It does not include hearing the caller, deciding that the caller has finished speaking, or turning text into audio. Microsoft's streaming transcriber reports partial text in just over 100 ms, per its announcement, and speech output adds its own start-up time.
So the delay a caller feels is a sum of several stages, each with its own tail. A slow turn is more likely to come from one stage being slow than from all of them at once.
What to do with the tail
Design for the slow turn, not the median.
- Test with your own prompts and record the median and the 95th percentile separately, from input received to first audio out.
- Cover the gap with a short acknowledgement phrase on turns that look slow, so silence does not read as a dropped call.
- Keep a timeout and a fallback response, so a stalled model call does not leave the line dead.
- Log slow turns with their prompt length, since long system prompts are a common cause.
Where non-live work goes
Not every audio task has a caller waiting. Narration, captions and music for a finished video can run as jobs, with a tail that nobody hears. A Sume job is read by polling or webhook, and its synchronous wait is capped at 30 seconds, per Sume jobs and results. Keep the strict latency work for the live path and let batch audio take the time it needs.
Remember that Mercury Voice is enterprise-only at launch, so the first step for most teams is a conversation with sales, covered in what to ask.
Sources
Related posts
More in Models
- MiniMax H3 2K and 4K upscale on Sume: H3 accepts them, H3 Max does not
minimax-h3 will price a 2K or 4K request even though its resolution list shows only 480p and 768p. minimax-h3-max rejects both. Costs for 5 to 15 seconds.
- MiniMax H3 or H3 Max on Sume: which id for a 5 to 15 second clip?
Choose minimax-h3 for a 480p or 768p draft at $0.075 a second, minimax-h3-max when you need 1080p. Same 5-15 s range, same references; Max costs more at 768p.
- MiniMax H3 reference-to-video: five images free, then $0.08 each
On the minimax-h3 id, reference-to-video adds a list fee for every reference image after the fifth. minimax-h3-max does not. Worked 10-second totals.
- MiniMax-H3 Turbo LoRA: 4 to 8 steps, Apache 2.0, base licence
The MiniMax-H3 Turbo LoRA cuts sampling to 4 to 8 steps and lists Apache 2.0, but the 33B base keeps its own terms. What it needs and the hosted route.
Written by Sume