AI voice agent cost per minute: speech-to-text, LLM and TTS stack
A dated cost sheet for the October 2026 voice stack: MAI-Transcribe-2-Streaming, Mercury Voice, MAI-Voice-2.1-Flash, and Sume's file STT and TTS, per minute.

A voice agent built from separate parts costs roughly three cents per minute of conversation at this week's list prices, by my arithmetic below: $0.009 for streaming transcription, about $0.009 for the language model at Inception's launch rate, and about $0.0135 for text to speech on MAI-Voice-2.1-Flash. That total depends on one assumption I state in the table and you should replace with your own: how many characters the agent speaks per minute.
Sume's rates for the same two speech steps are different in kind: a minute of file transcription is $0.01, and TTS is about $0.0475 per 1,000 characters, both as asynchronous jobs. They do not replace a live pipeline, but they are the right price for call recordings and pre-rendered lines.
What is each stage's price on the vendor's page?
Prices are as listed on each vendor page on 2026-10-03. MAI-Transcribe-2-Streaming is an introductory price through the end of the year, and the Mercury Voice price is a launch promotion. One figure, the batch MAI-Transcribe-2 price, is reported by a news site rather than Microsoft's page, and is not used in the totals.
| Stage | Product | Listed price | Per minute |
|---|---|---|---|
| Speech to text | MAI-Transcribe-2-Streaming | $0.54 per hour of audio (introductory, through year end) | $0.009 |
| Language model | Mercury Voice | About $0.009 per conversation minute (launch promotion, per Inception) | $0.009 |
| Text to speech | MAI-Voice-2.1-Flash | $15 per 1M characters | $0.0135 at 900 characters a minute (assumption) |
| Text to speech | MAI-Voice-2.1 | $22 per 1M characters | $0.0198 at 900 characters a minute (assumption) |
How do you recompute it for your own call?
The only inputs are minutes of caller audio, minutes of agent audio, and characters spoken. Here is the whole sheet as code, so you can swap in your numbers.
STT_PER_HOUR = 0.54 # MAI-Transcribe-2-Streaming, introductory
LLM_PER_MIN = 0.009 # Mercury Voice, Inception's launch figure
TTS_PER_M_CHARS = 15.00 # MAI-Voice-2.1-Flash
def call_cost(caller_minutes, call_minutes, agent_chars):
stt = caller_minutes * STT_PER_HOUR / 60
llm = call_minutes * LLM_PER_MIN
tts = agent_chars * TTS_PER_M_CHARS / 1e6
return round(stt + llm + tts, 4)
# Assumption: a 5-minute call, caller talks 2.5 minutes, agent says 2,250 characters
print(call_cost(2.5, 5, 2250))
What would the same speech cost on Sume?
Sume bills speech-to-text at $0.01 per audio minute and text to speech by character, so the file-based version of one call looks like this. This is the price for processing a recording or rendering fixed lines, not for a live turn.
| Item | Sume rate | Five-minute call |
|---|---|---|
| Transcribe the whole recording | $0.01 per audio minute | $0.05 |
| Render 2,250 characters of agent speech | About $0.0475 per 1,000 characters, rounded up to cents per job | About $0.107 before rounding, billed as $0.11 |
Which line moves the total the most?
In the sheet above the three live stages land within half a cent of each other, so no single vendor swap changes the answer much on its own. What changes it is behaviour: how long the agent talks, how much of each minute carries caller audio, and how many turns re-send history to the model.
Text to speech scales with characters, so a verbose agent costs more than a terse one at the same price per million characters. Speech to text scales with audio time, including silence, so a call with long pauses costs the same as a talkative one. The language model scales with turns and prompt size. Shortening the agent's answers is therefore the cheapest lever, because it lowers the synthesis bill, the model's output tokens and the call length together.
If you need to compare a different stack, change the three constants in the snippet and rerun it. For all-in-one realtime products priced per hour, realtime voice API cost per hour has the comparison. The sample call in the snippet, 2.5 minutes of caller audio in a five-minute call with 2,250 characters of agent speech, comes to about $0.10: $0.0225 for transcription, $0.045 for the model and $0.03375 for synthesis. It is an illustration with stated assumptions, not a measured call.
What should you not take from these totals?
For a wider comparison of all-in-one products, see Realtime voice API cost per hour, and for the speech steps alone, cheapest speech-to-text API per hour.
- The streaming transcription price counts audio you send, including silence, so a caller who pauses still costs the same.
- A launch promotion can change; Inception's page lists both standard and promotional rates.
- Your language-model bill can swing widely with history length, since the model re-reads the conversation each turn.
- None of these lists include telephony, hosting or your own engineering time.
Sources
Related posts
More in Pricing
- Change a person in a video with AI: Sume Recast price, not free
Changing a person in a video with Sume's H3 Max Recast is paid per source second: $0.375 at 768p, $0.5625 at 1080p. Costs for 5 to 30 second clips.
- Cost per row of a Sume bulk run: add up debited_usd_micros
Read usage.debited_usd_micros on each child receipt, wait for final to be true, and treat null as unknown. Why billable_amount alone understates a bulk run.
- Cost to remove backgrounds and upscale 500 product photos by API
Sume bills background removal at $0.0225 an image and upscale at $0.20, both flat. 500 photos through both steps cost $111.25; the table and a submit loop.
- Does AI video API billing round up? Sume's ceil rules per endpoint
A 5.2 s clip is billed as 6 s. Which Sume endpoints round seconds or minutes up, which prorate, and worked examples from the published rates.
Written by Sume