Can gpt-live-transcribe take an uploaded file? Endpoints vs Sume STT

OpenAI lists gpt-live-transcribe at $0.017/min and only for the realtime transcription endpoint. For a file, Sume STT is $0.01/min, up to 10 minutes per job.

5 min readSume
All posts

Not through its own model page's endpoint list. OpenAI's gpt-live-transcribe page lists one supported endpoint, realtime transcription at v1/realtime/transcription_sessions, and marks Chat Completions, Live sessions, Batch and the traditional transcription endpoints as unsupported. It is priced at $0.017 per minute of realtime audio. If you hold a finished file and want text, a file-based service fits better: Sume STT takes a public HTTPS audio_url and bills $0.01 per audio minute, with up to 600 seconds per job.

What the model page lists

The page describes a low-latency speech-to-text model for realtime use with tunable latency, unstructured context, keyword hints and multiple language hints. Input is audio and text, output is text. Rate limits start at 500 requests and 60,000 tokens per minute at Tier 1 and reach 10,000 requests and 780,000 tokens per minute at Tier 5. Those suit live sessions, where a connection stays open while someone speaks.

gpt-live-transcribe and Sume STT, vendor page and Sume docs, read 2026-10-05
Itemgpt-live-transcribeSume STT 1.0
Endpointv1/realtime/transcription_sessions onlyPOST /v1/stt-1.0/transcribe
Price$0.017 per minute$0.01 per audio minute
10 minutes of audio$0.17$0.10
60 minutes of audio$1.02$0.60 (six 10-minute jobs)
InputStreamed audio sessionPublic HTTPS audio_url, 10 minutes max

The cost of a one-hour recording

At $0.017 a minute, an hour is 60 x 0.017 = $1.02. At $0.01 a minute it is $0.60. The difference is 42 cents per hour of audio, or $42 across a hundred hours. Sume needs the hour split into six jobs of at most 10 minutes each, so the real cost of the saving is a splitting step. You can cut the file into ranges with Sume's audio detach tool and submit the pieces in parallel.

import os, requests
r = requests.post("https://api.sume.com/v1/stt-1.0/transcribe",
    headers={"x-api-key": os.environ["SUME_API_KEY"]},
    json={"audio_url": "https://example.com/call-part-1.mp3",
          "duration_seconds": 600,
          "segmentation": {"mode": "sentence"}}, timeout=60)
d = r.json()["data"]
print(d["job"]["status"], d["status_url"])

Which one to pick

Pick the realtime model when a person is speaking now and the text must follow within a second. Pick file transcription when the audio already exists and you want word timings back with the result. Sume always returns word timings and adds sentence segments if you ask for segmentation, which is handy for captions. Sume does not stream, so it does not replace a live caption feed.

Set duration_seconds when you know it: omitting it reserves one minute of usage.

Checking a model page before you build

Always read the supported-endpoints list on a model page, not just the price. The gpt-live-transcribe page lists its one supported endpoint and names the unsupported ones. A model that is cheap for the wrong shape of work is expensive, because you build a streaming client to push a finished file through it. Match the shape first: stream for live speech, file for recordings.

Splitting a long recording

A 60-minute meeting becomes six ranges of 10 minutes. Cut on silence when you can, so a sentence is not split across two jobs. Submit all six in parallel, then join the sentence segments in order using the offset of each range. Six jobs at $0.01 per minute is the same $0.60 as one hour at the minute rate, so splitting costs nothing extra beyond the code that does it.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume