Estimating a gpt-realtime-2.1 call from $32 / $64 per M audio tokens

gpt-realtime-2.1 lists audio input at $32, cached input at $0.40 and audio output at $64 per million tokens. Here is the formula and a worked example.

4 min readSume
All posts

gpt-realtime-2.1 charges per token, not per minute: $32 per million audio input tokens, $0.40 per million cached audio input tokens, and $64 per million audio output tokens. To estimate a call you need the number of tokens each side used, and the OpenAI pages I read do not convert minutes to tokens. So the honest method is to run a short test call, read the usage it reports, and scale. A worked example with assumed token counts is below, flagged as an assumption.

The listed prices

From OpenAI's gpt-realtime-2.1 model page, read 2026-10-05. The model is a reasoning speech-to-speech model available only through the Realtime API.

gpt-realtime-2.1 per 1M tokens and limits, OpenAI model page read 2026-10-05
ItemValue
Text input$4 per 1M tokens
Audio input$32 per 1M tokens
Cached audio input$0.40 per 1M tokens
Audio output$64 per 1M tokens
Context window128,000 tokens
Max output32,000 tokens
Tier 1 rate limit200 requests per minute
InterfaceRealtime API only

The formula

Call cost = (audio input tokens x $32 + cached input tokens x $0.40 + audio output tokens x $64 + text tokens x $4) / 1,000,000. Per thousand tokens that is $0.032 for audio in, $0.0004 for cached audio in, and $0.064 for audio out. Output costs twice as much as input, so a talkative agent is dearer than a quiet listener.

A worked example with assumed tokens

Assume, for illustration only, a call that used 20,000 uncached audio input tokens and 10,000 audio output tokens. These counts are made up; replace them with the usage you measure. Then input costs 20,000 x $32 / 1,000,000 = $0.64 and output costs 10,000 x $64 / 1,000,000 = $0.64, for $1.28 in total.

Context is where cost grows. In a long call each turn re-reads the conversation so far. Cached audio input at $0.40 per million is one-eightieth of the $32 uncached price (32 / 0.4 = 80), so cache hits matter a great deal on long sessions, though how much of your audio is cached depends on your integration and I do not state a rate here.

  • Measure a representative call end to end before extrapolating.
  • Count the output tokens at $64, not $32; they are the larger line in a chatty agent.
  • Watch the 128,000-token context window on long sessions.

Compared with Gemini Live's per-minute list

Google lists Gemini 3.8 Live audio input at $3.00 per million tokens, which its page shows as $0.005 a minute, and audio output at $12.00 per million, shown as $0.018 a minute. Because Google gives a per-minute equivalent and OpenAI does not, the two are not comparable until you measure tokens per minute on OpenAI's model. Compare after a test call, on your own audio, and including the model quality you need.

Where Sume fits

Sume does not offer a realtime voice-agent model. Its audio products are asynchronous: text-to-speech through Cartesia Sonic, speech-to-text on video clips, music and lip-sync clips. If you add a Sume job to a voice product, such as producing a recorded ad read, remember that Sume's pricing policy is provider list x 1.25 plus a 5.5% platform fee on customer spend.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume