Voice agent cost per minute: Gemini 3.8 Live vs gpt-realtime-2.1
Gemini 3.8 Live lists $0.005 per minute in and $0.018 out; gpt-realtime-2.1 lists $32 / $64 per 1M audio tokens. Convert tokens to minutes before you compare.

Gemini 3.8 Live is priced per minute: $0.005 per minute of audio in and $0.018 per minute out, which is $0.023 for a minute of audio in each direction. OpenAI prices gpt-realtime-2.1 per token: $32 per 1M audio input tokens and $64 per 1M audio output tokens. OpenAI's page, as I read it, does not give tokens per minute, so a per-minute figure needs an assumption. With the assumption below the OpenAI model comes out more than six times the Gemini price, but measure your own token usage before you rely on that.
Listed prices
| Item | Gemini 3.8 Live | gpt-realtime-2.1 |
|---|---|---|
| Audio input | $3.00 per 1M tokens ($0.005 per minute) | $32 per 1M tokens ($0.40 cached) |
| Audio output | $12.00 per 1M tokens ($0.018 per minute) | $64 per 1M tokens |
| Text input | Not stated in the text I read | $4 per 1M tokens |
| Context | Not stated in the text I read | 128,000 tokens; 32,000 max output |
| Access | Gemini API | Realtime API only; Tier 1 200 RPM |
Turning tokens into minutes
Assumption (mine, not OpenAI's): audio costs about 1,500 tokens a minute, the rate implied by Google's output line ($0.018 / $12 per 1M tokens = 1,500 tokens). OpenAI's tokenization may differ. At 1,500 tokens a minute: input is 1,500 x $32 / 1M = $0.048 and output is 1,500 x $64 / 1M = $0.096, so $0.144 for a minute each way. Against Gemini's $0.023, that is 0.144 / 0.023 = 6.3 times. If most of your input is cached at $0.40 per 1M, the input line falls to $0.0006 a minute and the gap narrows.
Real agents are not symmetrical. The model speaks perhaps half the time, and cached context, tool calls, text tokens and silence all change the bill. Take a usage report from a real session and compute cost per completed call.
What else belongs in the estimate
- Context growth: a long call re-sends history. gpt-realtime-2.1 has a 128,000-token context, so long calls can be expensive late in the call.
- Transport: telephony or WebRTC fees are outside both prices.
- Failure and retry rate: a call that must be redone doubles its cost.
- Quality: a cheaper minute that needs a human handoff more often costs more per resolved call.
Where Sume stands
Sume does not offer a live voice agent API or either of these models. Its audio products are asynchronous: TTS jobs on Cartesia Sonic, speech-to-text on hosted clips, music, and lip sync. These prices are a market reference only.
Sources
Related posts
More in Comparisons
- Google's Omni-first video rule has an expiry: Veo 3.1 ends Oct 22
Google's Gemini API docs name Omni Flash the default for video and keep Veo 3.1 for extension, last-frame control and legacy pipelines until October 22.
- Omni 1.1 Flash 1080p vs Wan 3.0 1080p: price for 30 seconds
A 30 second 1080p film costs $7.50 on Sume's Wan 3.0 in one request, or $5.625 as three 10 s Gemini Omni 1.1 Flash clips. What the cheaper route costs you.
- Omni 1.1 Flash extends to 40 s; Seedance 2.5 makes 30 s in one pass
Google extends Omni 1.1 Flash clips in 10 s steps to 40 s. Seedance 2.5 generates 30 s in one pass. How to get 40 s on Sume with Seedance and Timeline.
- Omni Flash 1.1 extends to 40 s: what Sume offers instead
Google says Omni Flash 1.1 extends in 10 s steps to 40 s. Sume lists Omni at 3-10 s, no extend. Four clips cost $5.00 at 720p; H3 Max needs three.
Written by Sume