MAI-Transcribe-2 speaker diarization vs Sume STT: no speaker field

MAI-Transcribe-2 attributes each segment to a distinct speaker. Sume STT has no diarize field: provider knobs are fixed server-side and you get words[].

4 min readSume
All posts

Microsoft says MAI-Transcribe-2 separates who said what, attributing each segment of audio to a distinct speaker. Sume STT gives you no speaker switch: the request schema states that provider knobs such as diarize and tag_audio_events are fixed server-side, and the documented output is words[] plus optional sentence segments.

Microsoft's claim is from its announcement post; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. This post does not say what the fixed server-side values are, because the docs do not.

What does Microsoft describe?

The post says multi-party audio, such as meetings, interviews, contact center calls and panel recordings, "comes back as structured, speaker-labeled dialogue instead of an undifferentiated block of text". It calls diarization a heavily requested capability. The announcement shows no request parameter or output field for speakers.

What does Sume expose instead?

The schema is closed: audio_url, optional segmentation, language_code, duration_seconds, metadata and the usual job mode fields. Nothing in it turns speaker separation on or off. The documented per-word entry is { word, start, end }, with no speaker property.

Speaker handling, Microsoft post vs Sume STT schema, read 2026-10-01.
NeedMAI-Transcribe-2 postSume `sume/stt-1.0`
Speaker-labeled segmentsYesNot documented
Caller diarize switchNot shownNone; knobs fixed server-side
Per-word timingYeswords[] with word, start, end
Sentence segmentsNot statedOptional segmentation.mode: "sentence"

Can I still separate speakers on Sume?

Not from the STT result. If you already know who speaks when, you can cut the audio into ranges per speaker and transcribe each piece, then label the output yourself. That costs one request per piece, and the duration_seconds hint tops out at 10 minutes. Detaching audio first gets you the right input shape.

When is a different service the right call?

If your product needs speaker-attributed transcripts out of the box, such as contact center review, pick a service that documents that output. If you need accurate word times for captions and cuts, the Sume words array covers that. The related Gemini diarization post draws the same line for another vendor.

Sources

Related posts

More in Models

All Models posts

Written by Sume