gpt-4o-transcribe-diarize retiring: Sume STT has no speaker labels

OpenAI lists gpt-4o-transcribe-diarize for removal on Feb 26, 2027. Sume STT returns words and sentences with timings, but no speaker field.

4 min readSume
All posts

OpenAI's deprecations page lists gpt-4o-transcribe-diarize for removal from its API on February 26, 2027, with gpt-live-transcribe or gpt-transcribe named as replacements. Sume STT (sume/stt-1.0) is not a drop-in answer to "who spoke when": it returns text, word timings and optional sentence segments, and no speaker field.

OpenAI facts are from its deprecations and speech-to-text pages, read 2026-09-30. Sume facts are from the OpenAPI document behind the API reference.

What does OpenAI say about the diarize model?

The speech-to-text guide says to use gpt-4o-transcribe-diarize only when you need to identify who speaks during different parts of a recording. You request the diarized_json response format to get segments with speaker, start and end, and for audio longer than 30 seconds you set chunking_strategy to "auto" or a voice activity detection configuration. The same removal date applies to the other OpenAI transcription models, and the guide also describes up to four known-speaker reference clips of 2 to 10 seconds that map segments to named speakers; Sume has no counterpart to that feature.

What does Sume STT return instead?

A completed job exposes text, language fields when available, and words[] with { word, start, end } in seconds from the audio start. If you send segmentation: { "mode": "sentence" }, Sume also groups the words into sentence segments derived from those timings. The request body is closed (additionalProperties: false), and the MCP tool description says diarize and tag_audio_events are fixed server-side, so sending them is rejected.

Speaker labelling, OpenAI diarize model versus Sume STT 1.0, read 2026-09-30.
NeedOpenAI diarize modelSume STT 1.0
Speaker label per segmentYes, in diarized_jsonNo speaker field
Start and end timesPer segmentPer word, plus optional sentence segments
Diarization option in requestResponse format and chunking_strategyRejected if sent
StatusRemoval Feb 26, 2027Current

What can I do if I need speaker turns on Sume?

Treat speaker attribution as a step outside STT. If each speaker was recorded on a separate file, transcribe each file as its own job and merge the results by word start time. If you only have a mixed recording, label turns from the sentence segments yourself or in a review pass. Sume does not promise automatic labels, so do not build a pipeline that expects them.

For the general interview flow, see how to transcribe an interview.

Should I migrate before February 2027?

If you rely on speaker labels today, check which replacement model OpenAI documents for that need before the removal date. If you only need text and timings, Sume STT covers that; see word timestamps.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume