Can gpt-live-transcribe take an uploaded file? Endpoints vs Sume STT
OpenAI lists gpt-live-transcribe at $0.017/min and only for the realtime transcription endpoint. For a file, Sume STT is $0.01/min, up to 10 minutes per job.

Not through its own model page's endpoint list. OpenAI's gpt-live-transcribe page lists one supported endpoint, realtime transcription at v1/realtime/transcription_sessions, and marks Chat Completions, Live sessions, Batch and the traditional transcription endpoints as unsupported. It is priced at $0.017 per minute of realtime audio. If you hold a finished file and want text, a file-based service fits better: Sume STT takes a public HTTPS audio_url and bills $0.01 per audio minute, with up to 600 seconds per job.
What the model page lists
The page describes a low-latency speech-to-text model for realtime use with tunable latency, unstructured context, keyword hints and multiple language hints. Input is audio and text, output is text. Rate limits start at 500 requests and 60,000 tokens per minute at Tier 1 and reach 10,000 requests and 780,000 tokens per minute at Tier 5. Those suit live sessions, where a connection stays open while someone speaks.
| Item | gpt-live-transcribe | Sume STT 1.0 |
|---|---|---|
| Endpoint | v1/realtime/transcription_sessions only | POST /v1/stt-1.0/transcribe |
| Price | $0.017 per minute | $0.01 per audio minute |
| 10 minutes of audio | $0.17 | $0.10 |
| 60 minutes of audio | $1.02 | $0.60 (six 10-minute jobs) |
| Input | Streamed audio session | Public HTTPS audio_url, 10 minutes max |
The cost of a one-hour recording
At $0.017 a minute, an hour is 60 x 0.017 = $1.02. At $0.01 a minute it is $0.60. The difference is 42 cents per hour of audio, or $42 across a hundred hours. Sume needs the hour split into six jobs of at most 10 minutes each, so the real cost of the saving is a splitting step. You can cut the file into ranges with Sume's audio detach tool and submit the pieces in parallel.
import os, requests
r = requests.post("https://api.sume.com/v1/stt-1.0/transcribe",
headers={"x-api-key": os.environ["SUME_API_KEY"]},
json={"audio_url": "https://example.com/call-part-1.mp3",
"duration_seconds": 600,
"segmentation": {"mode": "sentence"}}, timeout=60)
d = r.json()["data"]
print(d["job"]["status"], d["status_url"])Which one to pick
Pick the realtime model when a person is speaking now and the text must follow within a second. Pick file transcription when the audio already exists and you want word timings back with the result. Sume always returns word timings and adds sentence segments if you ask for segmentation, which is handy for captions. Sume does not stream, so it does not replace a live caption feed.
Set duration_seconds when you know it: omitting it reserves one minute of usage.
Checking a model page before you build
Always read the supported-endpoints list on a model page, not just the price. The gpt-live-transcribe page lists its one supported endpoint and names the unsupported ones. A model that is cheap for the wrong shape of work is expensive, because you build a streaming client to push a finished file through it. Match the shape first: stream for live speech, file for recordings.
Splitting a long recording
A 60-minute meeting becomes six ranges of 10 minutes. Cut on silence when you can, so a sentence is not split across two jobs. Submit all six in parallel, then join the sentence segments in order using the offset of each range. Six jobs at $0.01 per minute is the same $0.60 as one hour at the minute rate, so splitting costs nothing extra beyond the code that does it.
Sources
Related posts
More in Comparisons
- Cartesia Ink at $0.39 an hour vs ElevenLabs Scribe v2 at $0.22
Cartesia lists Ink STT at $0.39 an hour on the Scale plan; ElevenLabs lists Scribe v2 at $0.22. Ink costs 1.77x as much; Sume's STT is $0.01 a minute.
- Clean vs verbatim transcripts: MAI-Transcribe-2 style vs Sume captions
MAI-Transcribe-2 batch has transcribeStyle clean or verbatim. Sume STT has no style flag; for polished captions supply script_text and keep your wording.
- Clef or Clef-flash in front of a Sume agent: which tier to use
Cloudflare lists a median 38.8 ms for Clef-flash and 209.3 ms for Clef. Use the fast tier to gate runs and the larger one for the choices that cost real money.
- Cloudflare Workers AI TTS: Aura-1, Aura-2, MeloTTS vs Sume TTS
Cloudflare lists Aura-1 at $0.015 and Aura-2 at $0.030 per 1,000 characters and MeloTTS at $0.0002 per minute. Sume TTS 1.0 is $0.0475. Costs at three volumes.
Written by Sume