Mercury Voice vs text to speech: what a voice agent LLM does not do
Inception's Mercury Voice writes replies for voice agents. It does not synthesize audio. Where TTS and STT jobs sit around it, with Sume rates for each.

Mercury Voice is a diffusion language model that writes the reply in a voice agent, not a text-to-speech model. Inception reports a median time to first answer token under 320 ms, a P95 of 750 ms, and prices of $0.40 per million input tokens and $1.50 per million output tokens, halved during a launch discount with no stated end date (Inception Labs, read 2026-10-04). It reached general availability for enterprise customers on October 1, 2026, through an OpenAI-compatible endpoint. A speech layer still has to turn the reply into sound.
The three parts of a voice agent
| Stage | Job | Example |
|---|---|---|
| Hear | Speech to text | MAI-Transcribe-2-Streaming, streaming at about $0.54 per hour |
| Think | Write the reply | Mercury Voice, a language model |
| Speak | Text to speech | MAI-Voice-2.1-Flash at $15 per million characters |
Where Sume's audio jobs fit
Sume does not run a live voice agent loop. Its TTS and STT are job APIs, with a sync wait of at most 30 seconds (API reference). Two parts of an agent project still suit them.
- Fixed prompts: greetings, hold messages and error lines. Render each once, at $0.0475 per 1,000 characters, and play the file. The line bank guide shows the loop.
- Call review: transcribe a recorded call afterward at $0.01 per audio minute, up to 10 minutes per job. There are no speaker labels, so split the channels yourself first.
Be careful with the cost comparison
Mercury's price is per token and Microsoft's voice prices are per character, so there is no direct conversion. Inception states about $0.009 per minute of conversation for Mercury alone. That figure leaves out hearing and speaking. Add the speech-to-text and text-to-speech charges of whichever vendors you pick before you compare the total to anything else.
The launch discount is also a moving part. Ask whether a price quote includes it.
What to take from the launch
Language models for voice are getting cheap and fast, and the weak link moves to the audio layer. Decide early whether you need live speech, which is a streaming product, or pre-rendered speech, which a job API does well. Most teams need both, for different lines.
Sources
Related posts
More in Models
- MiniMax H3 Max: the prompt-adherence variant on Sume
fal describes MiniMax H3 Max as tuned for prompt adherence. On Sume, minimax-h3-max runs 480p to 1080p for 5 to 15 s with frames and references.
- MiniMax H3 limits: 9 images, 3 videos, 3 audio, file caps
MiniMax's H3 guide caps prompts at 7,000 characters and references at 9 images, 3 videos and 3 audio files. Cheat sheet with the Sume limits beside it.
- MiniMax H3 references on Sume: 9 images, 3 videos, 3 audio, 12 total
minimax-h3 and minimax-h3-max accept 9 images, 3 videos and 3 audio files, 12 in total, and audio cannot be the only reference. Duration rules and errors.
- Sume motion routes by length: Kling, Recast and Genjutsu
Three Sume routes drive a result from a source clip, each with a length rule: Kling takes duration_seconds, Recast follows the source, Genjutsu a range.
Written by Sume