Mercury Voice pricing: tokens to dollars per minute of talk
Mercury Voice lists at $0.40/$1.50 per million tokens, half off at launch, about $0.009 a minute. How to turn that into your bill, plus the speech steps.

Inception lists Mercury Voice at $0.40 per million input tokens and $1.50 per million output tokens, with a launch promotion at half that ($0.20 and $0.75), and says that works out to roughly $0.009 per minute of conversation. That is the price of the language model slot only. Speech in and speech out are billed separately.
The promotion is a launch discount, and the model is described as generally available for enterprise customers, with access through the Inception API as an OpenAI-compatible endpoint. Sume does not resell Mercury Voice, so the second half of this post is about pricing the speech around it with Sume's STT and TTS rates.
What does Inception's page say exactly?
The numbers below are copied from the Mercury Voice post, read 2026-10-03. The blog index lists the post under Sep 29, 2026, while the post itself states Oct 1, 2026.
| Item | Value on the page |
|---|---|
| Standard input | $0.40 per million tokens |
| Standard output | $1.50 per million tokens |
| Launch promotion (50% off) | $0.20 input, $0.75 output per million tokens |
| Per conversation minute | About $0.009, as stated by Inception |
| Context | 128K tokens, up to 50K output tokens |
| Reasoning settings | Low, medium, high |
| Availability | Generally available for enterprise customers; OpenAI-compatible endpoint |
How do you turn tokens into a per-minute figure?
Inception's $0.009 is a headline for its own test conversations. Your bill depends on how many tokens each turn sends (system prompt, history, tool output) and how many it returns. The formula is plain; the token counts must come from your own logs. The numbers in the snippet are hypothetical, there only to show the arithmetic at the promotional rates.
IN_PER_M, OUT_PER_M = 0.20, 0.75 # launch promotion, USD per million tokens
def turn_cost(input_tokens, output_tokens):
return input_tokens * IN_PER_M / 1e6 + output_tokens * OUT_PER_M / 1e6
# Hypothetical turn: 6,000 tokens of prompt and history, 300 tokens answered
per_turn = turn_cost(6000, 300)
print(round(per_turn, 6)) # 0.001425
print(round(per_turn * 6, 5)) # six such turns in one minute
What does the speech around it cost on Sume?
Sume STT bills $0.01 per audio minute (the reservation is one minute unless you send duration_seconds). TTS is character-priced: 38 micro-dollars of list per character, times 1.25, rounded up to cents per job, so about $0.0475 per 1,000 characters. If you assume an agent speaks 900 characters in a minute, which is an assumption and not a measured figure, that is about $0.043 of synthesis before cent rounding. Both are file jobs, so this applies to pre-rendered prompts and call recordings, not to a live turn.
| Step | Sume rate | Worked example |
|---|---|---|
| Speech to text | $0.01 per audio minute | 10-minute call recording: $0.10 |
| Text to speech | About $0.0475 per 1,000 characters, rounded up to cents per job | 900 characters: about $0.043 before rounding, billed as $0.05 |
| Language model | Not part of the Sume STT and TTS API | Use your own provider, such as Mercury Voice |
Why can the per-minute figure drift from the headline?
Inception's $0.009 is a stated approximation for a conversation minute, and it depends on a mix of input and output tokens that the page does not publish. Three things move your real figure. First, input tokens grow with every turn because the model re-reads the history, so a long call costs more per minute at the end than at the start. Second, tool calls add their results to the next prompt. Third, the reasoning setting (low, medium or high) changes how much the model writes before it answers, and output tokens cost almost four times as much as input at the listed rates ($1.50 against $0.40 standard, $0.75 against $0.20 promotional).
So the best practice is to log input and output token counts per turn from your first test calls and compute your own per-minute figure with the snippet above. Then look at the 128K context limit: a very long call with a large system prompt can approach it, and trimming history is both a cost and a safety measure.
What should you check before you compare totals?
The related realtime voice API cost per hour post compares all-in-one realtime products, which skip this three-bill arithmetic.
- Count all tokens the call sends, including history resent on every turn.
- Check whether the promotional price has an end date before you build a margin on it.
- Separate the live path (provider-streamed speech) from the recorded path (Sume jobs) when you add up cost.
- Ask Inception for the enterprise terms; the page says to contact sales and does not publish a self-serve signup.
Sources
Related posts
More in Pricing
- Balance needed for one AI generation call: published max per endpoint
Sume's catalog publishes estimated, minimum and maximum cents per endpoint. The worst-case hold for one call runs from 3 cents to $56.25 before a 402 can fire.
- Real-time AI avatar pricing: live minutes vs one render
Live avatars bill by conversation minute and concurrent stream; a rendered clip is paid once and watched by anyone. Tavus plan numbers and the break-even.
- Sume plans: Pro $40, Startup $120, Scale $400 and what they limit
What each Sume plan sets: monthly price, concurrent jobs, queue size and API write budget. Usage is billed at each model's rate, not by the plan.
- TTS cost by character: the same sentence in four languages
Sume bills TTS per character, not per second. One sentence counted in English, French, German and Korean, with the 1-cent floor and the 20,000-character cap.
Written by Sume