Mercury Voice is enterprise-only: pricing and what to ask sales

Mercury Voice is GA for enterprise customers only. List $0.40/$1.50 per million tokens, launch price $0.20/$0.75. Questions to ask before you commit.

5 min readSume
All posts

Can a solo developer use Mercury Voice today? Not directly: the Inception announcement says it is generally available for enterprise customers as of October 1, 2026, with access through the Inception API as an OpenAI API compatible endpoint, and it points interested teams to sales.

Standard rates are $0.40 per million input tokens and $1.50 per million output tokens. A 50 percent launch discount brings that to $0.20 and $0.75.

What the announcement states

Inception reports a median time to first answer token under 320 milliseconds and a p95 of 750 milliseconds. It puts a typical voice-agent conversation at about $0.009 per minute. Those are the vendor's figures, so treat them as claims to confirm on your own prompts.

Mercury Voice terms stated by Inception (read 2026-10-03)
ItemStated value
AvailabilityGenerally available, enterprise customers only, from October 1, 2026
InterfaceOpenAI API compatible endpoint through the Inception API
Standard price$0.40 per 1M input tokens, $1.50 per 1M output tokens
Launch price (50% off)$0.20 per 1M input tokens, $0.75 per 1M output tokens
Time to first answer tokenUnder 320 ms median, 750 ms p95
Per conversation minuteAbout $0.009 for typical voice-agent usage

Questions to ask sales

Because access is gated, the first call decides your plan. A short list makes it productive.

  • Does the launch discount have an end date, and what is the price after it ends?
  • What minimum spend or contract term applies to an enterprise account?
  • Do the latency numbers hold at your prompt length and tool-call pattern? The figures come from the vendor's own test set.
  • Which voice-agent platforms are supported in production? The page names LiveKit, Pipecat, Vapi and Retell as integration targets.
  • What are the rate limits and data-retention terms?

Mercury Voice is only one layer

Mercury Voice is the language-model step in a voice agent. It does not hear the caller and it does not speak. You still need speech-to-text in front and text-to-speech behind it, and the per-minute figure above does not include them. That is why the cost per minute of a full STT, LLM and TTS stack is a different number from the token price.

Sume sits in a different place in that picture. Its audio endpoints are job-based, so they suit the parts of the work that are not live: recorded narration, music, captions and transcripts of finished calls. A job returns by polling or by webhook, and synchronous waiting is capped at 30 seconds, as set out in Sume jobs and results. A live phone call needs a streaming path that Sume does not offer.

A reasonable next step

If you do not have an enterprise account, build and test the rest of the agent against any OpenAI-compatible model now, behind a single config value for base URL and model name. When Mercury Voice access arrives, you change that value and rerun your latency test instead of reworking the agent.

Sources

Related posts

More in Models

All Models posts

Written by Sume