gpt-live-transcribe has no word timestamps: what it means for captions
OpenAI's realtime guide says gpt-live-transcribe cannot return word-level timestamps. If you need karaoke captions, that decides which path to use.

Can OpenAI's gpt-live-transcribe time each word for captions? Not according to OpenAI's own guide. The Realtime transcription guide (read 2026-10-04) says gpt-live-transcribe returns transcript deltas as speech arrives and a final transcript when you commit each audio turn, and that it cannot provide word-level timestamps, speaker labels or confidence scores.
Word-level caption styles need word timing. So the choice between a live model and a file job is not only about price or latency; it decides whether per-word highlighting is possible at all.
What OpenAI documents
The guide names two models for realtime transcription and describes how they differ.
| Property | gpt-live-transcribe | gpt-transcribe |
|---|---|---|
| When text arrives | Deltas as speech arrives, final on commit | After a committed audio turn |
| Detected language in events | Not returned | Returned |
| Word timestamps | Not available | Not stated in the part of the guide we read |
| Turn detection settings | Omit them | Omit them |
What Sume's file job returns
Sume STT 1.0 is a file job. You pass a public HTTPS audio_url, optionally a language_code, and read text plus words[] where each item has word, start and end in seconds. Word timings are always returned, so there is no flag to switch on. Send segmentation: { mode: "sentence" } to also get gapless sentence segments.
That is what burned-in captions need. The video captions guide says a caption job with no script runs STT itself, and one that you give words skips it.
Which to choose
If you need text on screen while a person is still speaking, a live model is the right class, and Sume STT is not a streaming endpoint. If you are captioning a finished clip, a file job with word timing is the better fit, and it is simpler to retry.
Whichever you choose, record which path produced the timing. A live transcript with no word times cannot later be turned into karaoke highlighting; you would need to re-run the audio through a file job.
Sources
Related posts
More in Comparisons
- HeyGen translates 10 languages at once: what Sume does
HeyGen translate handles up to 10 target languages per job. Sume captions burn authored text per language, one $0.20 job per clip up to 60 seconds.
- HeyGen speaker rules as a checklist for a Sume face swap clip
HeyGen asks for a face within 45 degrees, one speaker at a time and under 10 feet. Use that as a pre-check, then Sume's own 4 to 15 second rules.
- Ideogram Plus or Pro vs API per image: what the page shows
Ideogram sells Plus at $15 and Pro at $42 a month in credits, and its page lists no per-image API price. How to compare it with a per-image row, with a script.
- Kling 4.0 voice reference vs Sume reference_audio_urls
Kling 4.0 Omni Reference accepts voice references. Sume takes 1-3 reference_audio_urls on models such as MiniMax H3; here is what each source documents.
Written by Sume