gpt-realtime-whisper $0.017 per minute vs Sume STT file jobs
OpenAI lists gpt-realtime-whisper at $0.017 a minute for streaming. Sume STT is $0.01 a minute but handles files, not live audio. Pick by workload.

OpenAI's model page lists gpt-realtime-whisper at $0.017 per minute, served through the realtime transcription sessions endpoint. Sume STT 1.0 costs $0.01 per audio minute but only transcribes a finished file at a public HTTPS URL. If you need words while someone is still speaking, Sume is not the tool; if you transcribe recordings, the lower per-minute rate and the job model apply.
What the OpenAI page lists
The page states the price, the endpoint and per-tier rate limits from 100 to 1,300 audio minutes per minute, which matters for live products with many simultaneous callers.
| Item | gpt-realtime-whisper | Sume STT 1.0 |
|---|---|---|
| Price | $0.017 per minute | $0.01 per audio minute |
| Mode | Streaming sessions | Async file job |
| Endpoint | v1/realtime/transcription_sessions | POST /v1/stt-1.0/transcribe |
| Input | Live audio | Public HTTPS audio_url |
The Sume side
A Sume STT request returns text and words[] with timings once the job completes. You poll GET /v1/jobs/:id/status or use a webhook, then read /result.
import os, requests
r = requests.post(
"https://api.sume.com/v1/stt-1.0/transcribe",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "stt-demo-001"},
json={"audio_url": "https://example.com/call.mp3",
"language_code": "en",
"duration_seconds": 300},
)
print(r.status_code, r.json())Decision rule
Use streaming for voice agents, live captions and call coaching. Use a file job for recordings, podcast archives, and the transcript step before captions or dubbing. Mixing both is normal: stream during the call, and send the saved recording for a cleaner final transcript.
Worked example
A support line with 100 hours of calls per month is a common size to test.
- 100 hours at $0.017 per minute: streaming transcription $102.00.
- The same 100 hours as recordings on Sume STT: $60.00.
- Streaming and file jobs are different products: only streaming returns words while the caller is still talking.
Checklist before you commit
Many teams use both: a live stream for agent assist and a recorded-file pass for the archive, which is where a file job fits.
- Write down your latency target before choosing.
- Check OpenAI's rate-limit tier for your account against peak concurrent calls.
- Keep call recordings only where your policies permit.
Sources
Related posts
More in Comparisons
- H3 Max Recast vs Genjutsu: which person swap to call on Sume
Sume lists two person-swap video rows. Recast: 1-4 people in a 5-30 s clip at 768p or 1080p. Genjutsu: 1-8 images at 480p or 720p. How to choose.
- Hedra 402 INSUFFICIENT_BALANCE vs Sume insufficient_credits
Hedra returns 402 INSUFFICIENT_BALANCE until you add funds; Sume returns 402 insufficient_credits. How each wallet check behaves and how to preflight a job.
- Hedra Avatar needs start frame and audio; Sume takes a handle and scri
Hedra Avatar generates from a start frame plus an audio track, up to 10 minutes. Sume's talking-video takes an avatar handle and a script, up to 60 seconds.
- Hedra Character 3 allows 10 minutes; what is the Sume equivalent?
Hedra lists Character 3 at up to 10 minutes. Sume's Avatar 1.0 takes 4-60 seconds per job, so 10 minutes means ten or more jobs joined on a timeline.
Written by Sume