gpt-live-transcribe has no word timestamps: what it means for captions

OpenAI's realtime guide says gpt-live-transcribe cannot return word-level timestamps. If you need karaoke captions, that decides which path to use.

5 min readSume
All posts

Can OpenAI's gpt-live-transcribe time each word for captions? Not according to OpenAI's own guide. The Realtime transcription guide (read 2026-10-04) says gpt-live-transcribe returns transcript deltas as speech arrives and a final transcript when you commit each audio turn, and that it cannot provide word-level timestamps, speaker labels or confidence scores.

Word-level caption styles need word timing. So the choice between a live model and a file job is not only about price or latency; it decides whether per-word highlighting is possible at all.

What OpenAI documents

The guide names two models for realtime transcription and describes how they differ.

gpt-live-transcribe and gpt-transcribe in the realtime guide, read 2026-10-04
Propertygpt-live-transcribegpt-transcribe
When text arrivesDeltas as speech arrives, final on commitAfter a committed audio turn
Detected language in eventsNot returnedReturned
Word timestampsNot availableNot stated in the part of the guide we read
Turn detection settingsOmit themOmit them

What Sume's file job returns

Sume STT 1.0 is a file job. You pass a public HTTPS audio_url, optionally a language_code, and read text plus words[] where each item has word, start and end in seconds. Word timings are always returned, so there is no flag to switch on. Send segmentation: { mode: "sentence" } to also get gapless sentence segments.

That is what burned-in captions need. The video captions guide says a caption job with no script runs STT itself, and one that you give words skips it.

Which to choose

If you need text on screen while a person is still speaking, a live model is the right class, and Sume STT is not a streaming endpoint. If you are captioning a finished clip, a file job with word timing is the better fit, and it is simpler to retry.

Whichever you choose, record which path produced the timing. A live transcript with no word times cannot later be turned into karaoke highlighting; you would need to re-run the audio through a file job.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume