gpt-4o-transcribe-diarize retiring: Sume STT has no speaker labels
OpenAI lists gpt-4o-transcribe-diarize for removal on Feb 26, 2027. Sume STT returns words and sentences with timings, but no speaker field.

OpenAI's deprecations page lists gpt-4o-transcribe-diarize for removal from its API on February 26, 2027, with gpt-live-transcribe or gpt-transcribe named as replacements. Sume STT (sume/stt-1.0) is not a drop-in answer to "who spoke when": it returns text, word timings and optional sentence segments, and no speaker field.
OpenAI facts are from its deprecations and speech-to-text pages, read 2026-09-30. Sume facts are from the OpenAPI document behind the API reference.
What does OpenAI say about the diarize model?
The speech-to-text guide says to use gpt-4o-transcribe-diarize only when you need to identify who speaks during different parts of a recording. You request the diarized_json response format to get segments with speaker, start and end, and for audio longer than 30 seconds you set chunking_strategy to "auto" or a voice activity detection configuration. The same removal date applies to the other OpenAI transcription models, and the guide also describes up to four known-speaker reference clips of 2 to 10 seconds that map segments to named speakers; Sume has no counterpart to that feature.
What does Sume STT return instead?
A completed job exposes text, language fields when available, and words[] with { word, start, end } in seconds from the audio start. If you send segmentation: { "mode": "sentence" }, Sume also groups the words into sentence segments derived from those timings. The request body is closed (additionalProperties: false), and the MCP tool description says diarize and tag_audio_events are fixed server-side, so sending them is rejected.
| Need | OpenAI diarize model | Sume STT 1.0 |
|---|---|---|
| Speaker label per segment | Yes, in diarized_json | No speaker field |
| Start and end times | Per segment | Per word, plus optional sentence segments |
| Diarization option in request | Response format and chunking_strategy | Rejected if sent |
| Status | Removal Feb 26, 2027 | Current |
What can I do if I need speaker turns on Sume?
Treat speaker attribution as a step outside STT. If each speaker was recorded on a separate file, transcribe each file as its own job and merge the results by word start time. If you only have a mixed recording, label turns from the sentence segments yourself or in a review pass. Sume does not promise automatic labels, so do not build a pipeline that expects them.
For the general interview flow, see how to transcribe an interview.
Should I migrate before February 2027?
If you rely on speaker labels today, check which replacement model OpenAI documents for that need before the removal date. If you only need text and timings, Sume STT covers that; see word timestamps.
Sources
Related posts
More in Developers
- GPT Image 2.5 1080x1350: why the size fails and what to send
1080x1350 breaks GPT Image 2.5's multiple-of-16 size rule. 1080 is not a multiple of 16. 1088x1360 (exact 4:5) meets the listed rules. Per OpenAI and Sume docs.
- GPT Image 2.5 Batch API: not supported, so fan out jobs
OpenAI's model page lists Batch as unsupported for GPT Image 2.5. On Sume, submit many async or webhook jobs to /v1/images and collect the results.
- GPT Image 2.5 mask edit: does the mask need an alpha channel?
OpenAI's API says the mask must contain an alpha channel. Sume's docs list a public HTTPS mask_url for GPT Image 2.5 edits but state no mask format rule.
- GPT Image 2.5 moderation low: can you send it through Sume?
OpenAI lists a moderation parameter (auto or low) for GPT Image 2.5. Sume's request table does not list it, so check the catalog before sending it.
Written by Sume