Voice agent cost per minute: Gemini 3.8 Live vs gpt-realtime-2.1

Gemini 3.8 Live lists $0.005 per minute in and $0.018 out; gpt-realtime-2.1 lists $32 / $64 per 1M audio tokens. Convert tokens to minutes before you compare.

4 min readSume
All posts

Gemini 3.8 Live is priced per minute: $0.005 per minute of audio in and $0.018 per minute out, which is $0.023 for a minute of audio in each direction. OpenAI prices gpt-realtime-2.1 per token: $32 per 1M audio input tokens and $64 per 1M audio output tokens. OpenAI's page, as I read it, does not give tokens per minute, so a per-minute figure needs an assumption. With the assumption below the OpenAI model comes out more than six times the Gemini price, but measure your own token usage before you rely on that.

Listed prices

Google pricing page and OpenAI model page, read 2026-10-05
ItemGemini 3.8 Livegpt-realtime-2.1
Audio input$3.00 per 1M tokens ($0.005 per minute)$32 per 1M tokens ($0.40 cached)
Audio output$12.00 per 1M tokens ($0.018 per minute)$64 per 1M tokens
Text inputNot stated in the text I read$4 per 1M tokens
ContextNot stated in the text I read128,000 tokens; 32,000 max output
AccessGemini APIRealtime API only; Tier 1 200 RPM

Turning tokens into minutes

Assumption (mine, not OpenAI's): audio costs about 1,500 tokens a minute, the rate implied by Google's output line ($0.018 / $12 per 1M tokens = 1,500 tokens). OpenAI's tokenization may differ. At 1,500 tokens a minute: input is 1,500 x $32 / 1M = $0.048 and output is 1,500 x $64 / 1M = $0.096, so $0.144 for a minute each way. Against Gemini's $0.023, that is 0.144 / 0.023 = 6.3 times. If most of your input is cached at $0.40 per 1M, the input line falls to $0.0006 a minute and the gap narrows.

Real agents are not symmetrical. The model speaks perhaps half the time, and cached context, tool calls, text tokens and silence all change the bill. Take a usage report from a real session and compute cost per completed call.

What else belongs in the estimate

  • Context growth: a long call re-sends history. gpt-realtime-2.1 has a 128,000-token context, so long calls can be expensive late in the call.
  • Transport: telephony or WebRTC fees are outside both prices.
  • Failure and retry rate: a call that must be redone doubles its cost.
  • Quality: a cheaper minute that needs a human handoff more often costs more per resolved call.

Where Sume stands

Sume does not offer a live voice agent API or either of these models. Its audio products are asynchronous: TTS jobs on Cartesia Sonic, speech-to-text on hosted clips, music, and lip sync. These prices are a market reference only.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume