gpt-live-transcribe: realtime STT at $0.017 a minute
gpt-live-transcribe costs $0.017 a minute ($1.02 an hour) and runs only on the Realtime transcription sessions endpoint. Use gpt-transcribe for files.

gpt-live-transcribe is OpenAI's streaming speech-to-text model, priced at $0.017 per minute, which is $1.02 per hour. It works only on the v1/realtime/transcription_sessions endpoint. If your audio is a finished recording, OpenAI's own guide points you to gpt-transcribe at $0.0045 a minute instead, which is about 3.8 times cheaper.
What the model page says
Facts from OpenAI's model page and transcription guide, read 2026-10-05.
| Item | Detail |
|---|---|
| Purpose | Streaming speech-to-text |
| Price | $0.017 per minute |
| Endpoint | Only v1/realtime/transcription_sessions |
| Latency | Tunable |
| Accuracy aids | Keyword hints and multiple language hints |
| For completed recordings | Use gpt-transcribe instead |
How $0.017 compares with other live options
Streaming prices vary by vendor and by billing basis. AssemblyAI's page notes that streaming is billed on session duration, not audio duration, so a session left open costs money even in silence. The OpenAI listing says per minute; I could not tell from the page whether it uses the same basis, so check before you hold a session open for hours.
| Service | Listed price | Per hour |
|---|---|---|
| AssemblyAI Universal-Streaming | $0.15 per hour | $0.15 |
| ElevenLabs Scribe v2 Realtime | $0.39 per hour (about 150 ms) | $0.39 |
| AssemblyAI Universal-3.6 Pro Realtime | $0.45 per hour | $0.45 |
| OpenAI gpt-live-transcribe | $0.017 per minute | $1.02 |
| OpenAI gpt-realtime-translate | $0.034 per minute | $2.04 |
When the premium is worth thinking about
On listed numbers, gpt-live-transcribe is the most expensive transcription row above (the translate row is a different product): $1.02 an hour against $0.39 for Scribe v2 Realtime, which is 2.6 times as much (1.02 / 0.39). The price does not say what you get for the difference. The model page describes tunable latency and keyword and language hints, which matter for domain terms and mixed-language speech, but I did not read accuracy or latency numbers for any of these services, so run your own audio through two of them before you pick.
- Use gpt-live-transcribe when you are already inside OpenAI's Realtime stack and want hints for product names.
- Use a file endpoint, such as gpt-transcribe, for anything you can wait to finish.
- OpenAI's guide says speaker labels need gpt-4o-transcribe-diarize, and timestamps or subtitle formats need whisper-1, so a live transcript alone may not give you either.
What Sume offers
Sume does not transcribe live streams. Its speech-to-text is a flag on Video inspect: transcribe: true on one hosted clip, at $0.01 per audio minute, with an optional language_code hint and sentence segments. For a meeting you record and upload afterwards, that works; for captions on a live call, use a streaming vendor.
Sources
Related posts
More in Models
- xAI recommends Grok Imagine Video 1.5 ($0.08/s) over the $0.05 model
xAI lists grok-imagine-video-1.5 at $0.080 per second and grok-imagine-video at $0.050, and recommends 1.5. Sume carries only 1.5, image-to-video.
- grok-imagine-video-1.5 on Sume: needs an image, seven fields refused
grok-imagine-video-1.5 is image-to-video only on Sume: send one image, no end frame, no reference video or audio, no aspect_ratio, no generate_audio.
- H3 Max 1080p regenerates from 768p: the price of the extra step
MiniMax says its 2K path regenerates in context. Sume documents H3 Max 1080p as a latent refinement of 768p, $0.20 vs $0.10 per second. When the doubling pays.
- H3 or H3 Max? Pick the Sume row by job, with per-clip prices
Both read references and make stereo audio. H3 is $0.075 per second at 768p, H3 Max $0.10 with a 1080p row. A job-by-job guide with 10-second prices.
Written by Sume