Realtime voice API cost per hour: Grok, GPT-Live-1, Gemini Live
Per-hour cost of realtime voice APIs from the vendors' own pages, set against Sume's async STT plus TTS jobs, with the arithmetic and its assumptions shown.

A realtime voice session costs roughly $1.38 to $4.80 per hour of continuous audio, depending on the vendor, based on the per-minute rates each vendor lists today. Sume does not run live sessions, so it is not a like-for-like replacement. If your product can tolerate a few seconds per turn, the same hour built from Sume's async speech-to-text and text-to-speech jobs comes to about $3.17 under the assumptions below.
Vendor rates, read 1 October 2026
Each row is the vendor's published rate, converted to an hourly figure by multiplying by 60. A session that is idle or only half talking costs less where the vendor bills by audio duration.
| Model | Published rate | Per hour |
|---|---|---|
| Grok Voice Think Fast 2.0 (speech to speech) | $0.08 per minute | $4.80 |
| OpenAI gpt-live-1 | $0.05 per minute, billed per second | $3.00 |
| Gemini 3.8 Live, audio in | $0.005 per minute | $0.30 |
| Gemini 3.8 Live, audio out | $0.018 per minute | $1.08 |
The Sume pipeline, with the arithmetic
Sume prices are listed on the API pricing page: speech-to-text is $0.01 per audio minute and text-to-speech is $0.0475 per 1,000 characters. To compare an hour, assume the user speaks for 60 minutes and the agent speaks for 60 minutes at about 150 words per minute and 6 characters per word including spaces. That is 54,000 characters of speech, which is an assumption, not a measurement.
| Step | Quantity | Cost |
|---|---|---|
| Transcribe user audio | 60 audio minutes at $0.01 | $0.60 |
| Synthesize agent reply | 54 x 1,000 characters at $0.0475 | $2.565 |
| Total | about $3.17 |
Limits that matter more than price
The two shapes behave differently, and cost is rarely the deciding factor.
- Sume's speech-to-text takes a public HTTPS audio URL and handles up to 10 minutes per request, so you transcribe finished clips, not an open microphone stream.
- Text-to-speech is a job: submit, wait for the job, fetch the result. See jobs and results.
- Sume does not bundle a language model, so the reply text has to come from your own step. Realtime models include that reasoning in the per-minute price; the Sume total above does not.
- Vendor rates change. Recheck the three pages before you budget.
When Sume fits
Use async jobs for narration, voicemail summaries, recorded call review and voice notes, where a turn can take seconds. Use a realtime session when someone is waiting on the line. Related reading: realtime model versus async narration and which Sume audio endpoint to call.
Sources
More in Pricing
- Recraft API pricing: 1,000 API units per dollar, V4.1 Flash $0.007
Recraft sells prepaid API units at $1.00 for 1,000. V4.1 Flash is 7 units ($0.007) per raster image, V4.1 is 35 units ($0.035), V4.1 Pro is 210 units ($0.21).
- Recraft vector price: $0.08 per SVG image vs $0.04 raster
On Recraft's price page, V3 Vector is $0.08 per image against $0.04 for V3 raster, and vectorizing an existing image is $0.01 per request. See the full table.
- Reference ingest cost: free manifest, STT only if speech is heard
Sume reference ingest reads a clip of up to 300 seconds into shots, OCR and audio facts for free; transcription bills $0.01 a minute only if speech is detected.
- Reference ingest transcript: allow_billed_stt and skipped STT
Reference ingest only bills a transcript when you opt in and the clip has speech. What allow_billed_stt reserves and what the stt_skipped warnings mean.
Written by Sume