Realtime voice API cost per hour: Grok, GPT-Live-1, Gemini Live

Per-hour cost of realtime voice APIs from the vendors' own pages, set against Sume's async STT plus TTS jobs, with the arithmetic and its assumptions shown.

5 min readSume
All posts

A realtime voice session costs roughly $1.38 to $4.80 per hour of continuous audio, depending on the vendor, based on the per-minute rates each vendor lists today. Sume does not run live sessions, so it is not a like-for-like replacement. If your product can tolerate a few seconds per turn, the same hour built from Sume's async speech-to-text and text-to-speech jobs comes to about $3.17 under the assumptions below.

Vendor rates, read 1 October 2026

Each row is the vendor's published rate, converted to an hourly figure by multiplying by 60. A session that is idle or only half talking costs less where the vendor bills by audio duration.

Realtime voice pricing from vendor pages, read 2026-10-01. Hourly column is the per-minute rate times 60.
ModelPublished ratePer hour
Grok Voice Think Fast 2.0 (speech to speech)$0.08 per minute$4.80
OpenAI gpt-live-1$0.05 per minute, billed per second$3.00
Gemini 3.8 Live, audio in$0.005 per minute$0.30
Gemini 3.8 Live, audio out$0.018 per minute$1.08

The Sume pipeline, with the arithmetic

Sume prices are listed on the API pricing page: speech-to-text is $0.01 per audio minute and text-to-speech is $0.0475 per 1,000 characters. To compare an hour, assume the user speaks for 60 minutes and the agent speaks for 60 minutes at about 150 words per minute and 6 characters per word including spaces. That is 54,000 characters of speech, which is an assumption, not a measurement.

Sume async cost for one hour each way, using list prices from the Sume API pricing page, read 2026-10-01.
StepQuantityCost
Transcribe user audio60 audio minutes at $0.01$0.60
Synthesize agent reply54 x 1,000 characters at $0.0475$2.565
Totalabout $3.17

Limits that matter more than price

The two shapes behave differently, and cost is rarely the deciding factor.

  • Sume's speech-to-text takes a public HTTPS audio URL and handles up to 10 minutes per request, so you transcribe finished clips, not an open microphone stream.
  • Text-to-speech is a job: submit, wait for the job, fetch the result. See jobs and results.
  • Sume does not bundle a language model, so the reply text has to come from your own step. Realtime models include that reasoning in the per-minute price; the Sume total above does not.
  • Vendor rates change. Recheck the three pages before you budget.

When Sume fits

Use async jobs for narration, voicemail summaries, recorded call review and voice notes, where a turn can take seconds. Use a realtime session when someone is waiting on the line. Related reading: realtime model versus async narration and which Sume audio endpoint to call.

Sources

More in Pricing

All Pricing posts

Written by Sume