Mercury Voice vs text to speech: what a voice agent LLM does not do

Inception's Mercury Voice writes replies for voice agents. It does not synthesize audio. Where TTS and STT jobs sit around it, with Sume rates for each.

4 min readSume
All posts

Mercury Voice is a diffusion language model that writes the reply in a voice agent, not a text-to-speech model. Inception reports a median time to first answer token under 320 ms, a P95 of 750 ms, and prices of $0.40 per million input tokens and $1.50 per million output tokens, halved during a launch discount with no stated end date (Inception Labs, read 2026-10-04). It reached general availability for enterprise customers on October 1, 2026, through an OpenAI-compatible endpoint. A speech layer still has to turn the reply into sound.

The three parts of a voice agent

Voice agent stages and the products named in this post (read 2026-10-04)
StageJobExample
HearSpeech to textMAI-Transcribe-2-Streaming, streaming at about $0.54 per hour
ThinkWrite the replyMercury Voice, a language model
SpeakText to speechMAI-Voice-2.1-Flash at $15 per million characters

Where Sume's audio jobs fit

Sume does not run a live voice agent loop. Its TTS and STT are job APIs, with a sync wait of at most 30 seconds (API reference). Two parts of an agent project still suit them.

  • Fixed prompts: greetings, hold messages and error lines. Render each once, at $0.0475 per 1,000 characters, and play the file. The line bank guide shows the loop.
  • Call review: transcribe a recorded call afterward at $0.01 per audio minute, up to 10 minutes per job. There are no speaker labels, so split the channels yourself first.

Be careful with the cost comparison

Mercury's price is per token and Microsoft's voice prices are per character, so there is no direct conversion. Inception states about $0.009 per minute of conversation for Mercury alone. That figure leaves out hearing and speaking. Add the speech-to-text and text-to-speech charges of whichever vendors you pick before you compare the total to anything else.

The launch discount is also a moving part. Ask whether a price quote includes it.

What to take from the launch

Language models for voice are getting cheap and fast, and the weak link moves to the audio layer. Decide early whether you need live speech, which is a streaming product, or pre-rendered speech, which a job API does well. Most teams need both, for different lines.

Sources

Related posts

More in Models

All Models posts

Written by Sume