AI voice agent cost per minute: speech-to-text, LLM and TTS stack

A dated cost sheet for the October 2026 voice stack: MAI-Transcribe-2-Streaming, Mercury Voice, MAI-Voice-2.1-Flash, and Sume's file STT and TTS, per minute.

5 min readSume
All posts

A voice agent built from separate parts costs roughly three cents per minute of conversation at this week's list prices, by my arithmetic below: $0.009 for streaming transcription, about $0.009 for the language model at Inception's launch rate, and about $0.0135 for text to speech on MAI-Voice-2.1-Flash. That total depends on one assumption I state in the table and you should replace with your own: how many characters the agent speaks per minute.

Sume's rates for the same two speech steps are different in kind: a minute of file transcription is $0.01, and TTS is about $0.0475 per 1,000 characters, both as asynchronous jobs. They do not replace a live pipeline, but they are the right price for call recordings and pre-rendered lines.

What is each stage's price on the vendor's page?

Prices are as listed on each vendor page on 2026-10-03. MAI-Transcribe-2-Streaming is an introductory price through the end of the year, and the Mercury Voice price is a launch promotion. One figure, the batch MAI-Transcribe-2 price, is reported by a news site rather than Microsoft's page, and is not used in the totals.

Voice stack list prices read 2026-10-03 from Microsoft AI and Inception.
StageProductListed pricePer minute
Speech to textMAI-Transcribe-2-Streaming$0.54 per hour of audio (introductory, through year end)$0.009
Language modelMercury VoiceAbout $0.009 per conversation minute (launch promotion, per Inception)$0.009
Text to speechMAI-Voice-2.1-Flash$15 per 1M characters$0.0135 at 900 characters a minute (assumption)
Text to speechMAI-Voice-2.1$22 per 1M characters$0.0198 at 900 characters a minute (assumption)

How do you recompute it for your own call?

The only inputs are minutes of caller audio, minutes of agent audio, and characters spoken. Here is the whole sheet as code, so you can swap in your numbers.

STT_PER_HOUR = 0.54          # MAI-Transcribe-2-Streaming, introductory
LLM_PER_MIN = 0.009          # Mercury Voice, Inception's launch figure
TTS_PER_M_CHARS = 15.00      # MAI-Voice-2.1-Flash

def call_cost(caller_minutes, call_minutes, agent_chars):
    stt = caller_minutes * STT_PER_HOUR / 60
    llm = call_minutes * LLM_PER_MIN
    tts = agent_chars * TTS_PER_M_CHARS / 1e6
    return round(stt + llm + tts, 4)

# Assumption: a 5-minute call, caller talks 2.5 minutes, agent says 2,250 characters
print(call_cost(2.5, 5, 2250))

What would the same speech cost on Sume?

Sume bills speech-to-text at $0.01 per audio minute and text to speech by character, so the file-based version of one call looks like this. This is the price for processing a recording or rendering fixed lines, not for a live turn.

Sume file-job rates from the pricing code and OpenAPI document, read 2026-10-03; the 2,250-character figure is the assumption above.
ItemSume rateFive-minute call
Transcribe the whole recording$0.01 per audio minute$0.05
Render 2,250 characters of agent speechAbout $0.0475 per 1,000 characters, rounded up to cents per jobAbout $0.107 before rounding, billed as $0.11

Which line moves the total the most?

In the sheet above the three live stages land within half a cent of each other, so no single vendor swap changes the answer much on its own. What changes it is behaviour: how long the agent talks, how much of each minute carries caller audio, and how many turns re-send history to the model.

Text to speech scales with characters, so a verbose agent costs more than a terse one at the same price per million characters. Speech to text scales with audio time, including silence, so a call with long pauses costs the same as a talkative one. The language model scales with turns and prompt size. Shortening the agent's answers is therefore the cheapest lever, because it lowers the synthesis bill, the model's output tokens and the call length together.

If you need to compare a different stack, change the three constants in the snippet and rerun it. For all-in-one realtime products priced per hour, realtime voice API cost per hour has the comparison. The sample call in the snippet, 2.5 minutes of caller audio in a five-minute call with 2,250 characters of agent speech, comes to about $0.10: $0.0225 for transcription, $0.045 for the model and $0.03375 for synthesis. It is an illustration with stated assumptions, not a measured call.

What should you not take from these totals?

For a wider comparison of all-in-one products, see Realtime voice API cost per hour, and for the speech steps alone, cheapest speech-to-text API per hour.

  • The streaming transcription price counts audio you send, including silence, so a caller who pauses still costs the same.
  • A launch promotion can change; Inception's page lists both standard and promotional rates.
  • Your language-model bill can swing widely with history length, since the model re-reads the conversation each turn.
  • None of these lists include telephony, hosting or your own engineering time.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume