Estimating a gpt-realtime-2.1 call from $32 / $64 per M audio tokens
gpt-realtime-2.1 lists audio input at $32, cached input at $0.40 and audio output at $64 per million tokens. Here is the formula and a worked example.

gpt-realtime-2.1 charges per token, not per minute: $32 per million audio input tokens, $0.40 per million cached audio input tokens, and $64 per million audio output tokens. To estimate a call you need the number of tokens each side used, and the OpenAI pages I read do not convert minutes to tokens. So the honest method is to run a short test call, read the usage it reports, and scale. A worked example with assumed token counts is below, flagged as an assumption.
The listed prices
From OpenAI's gpt-realtime-2.1 model page, read 2026-10-05. The model is a reasoning speech-to-speech model available only through the Realtime API.
| Item | Value |
|---|---|
| Text input | $4 per 1M tokens |
| Audio input | $32 per 1M tokens |
| Cached audio input | $0.40 per 1M tokens |
| Audio output | $64 per 1M tokens |
| Context window | 128,000 tokens |
| Max output | 32,000 tokens |
| Tier 1 rate limit | 200 requests per minute |
| Interface | Realtime API only |
The formula
Call cost = (audio input tokens x $32 + cached input tokens x $0.40 + audio output tokens x $64 + text tokens x $4) / 1,000,000. Per thousand tokens that is $0.032 for audio in, $0.0004 for cached audio in, and $0.064 for audio out. Output costs twice as much as input, so a talkative agent is dearer than a quiet listener.
A worked example with assumed tokens
Assume, for illustration only, a call that used 20,000 uncached audio input tokens and 10,000 audio output tokens. These counts are made up; replace them with the usage you measure. Then input costs 20,000 x $32 / 1,000,000 = $0.64 and output costs 10,000 x $64 / 1,000,000 = $0.64, for $1.28 in total.
Context is where cost grows. In a long call each turn re-reads the conversation so far. Cached audio input at $0.40 per million is one-eightieth of the $32 uncached price (32 / 0.4 = 80), so cache hits matter a great deal on long sessions, though how much of your audio is cached depends on your integration and I do not state a rate here.
- Measure a representative call end to end before extrapolating.
- Count the output tokens at $64, not $32; they are the larger line in a chatty agent.
- Watch the 128,000-token context window on long sessions.
Compared with Gemini Live's per-minute list
Google lists Gemini 3.8 Live audio input at $3.00 per million tokens, which its page shows as $0.005 a minute, and audio output at $12.00 per million, shown as $0.018 a minute. Because Google gives a per-minute equivalent and OpenAI does not, the two are not comparable until you measure tokens per minute on OpenAI's model. Compare after a test call, on your own audio, and including the model quality you need.
Where Sume fits
Sume does not offer a realtime voice-agent model. Its audio products are asynchronous: text-to-speech through Cartesia Sonic, speech-to-text on video clips, music and lip-sync clips. If you add a Sume job to a voice product, such as producing a recorded ad read, remember that Sume's pricing policy is provider list x 1.25 plus a 5.5% platform fee on customer spend.
Sources
Related posts
More in Pricing
- gpt-transcribe at $0.0045 a minute vs Scribe v2 and Nova-3
Of the three, Scribe v2 lists lowest at $0.22 an hour, then Nova-3 mono at $0.258 and gpt-transcribe at $0.27. Per-hour table and what Sume bills.
- Grok Imagine image prices: $0.02, $0.04, $0.05 and Sume's one row
xAI lists three Grok image prices ($0.02, $0.04, $0.05). Sume lists one x-ai/grok-image row; read its live endpoints pricing array for the billed figure.
- Grok Imagine video cost: a 15-second clip is $1.20 at xAI's 1.5 rate
xAI's video guide says duration and resolution both set the cost, with 15 seconds the maximum. At $0.080 per second a 15 s clip is $1.20.
- H3 2K at $1.63 vs Omni 1080p at $1.88: 10-second price check
On Sume's price list, a 10-second H3 2K upscale is $1.63, Omni Flash 1.1 at 1080p is $1.88 and 4K is $3.75, H3 4K $2.00. What each row lists today.
Written by Sume