MAI-Transcribe-2 speaker diarization vs Sume STT: no speaker field
MAI-Transcribe-2 attributes each segment to a distinct speaker. Sume STT has no diarize field: provider knobs are fixed server-side and you get words[].

Microsoft says MAI-Transcribe-2 separates who said what, attributing each segment of audio to a distinct speaker. Sume STT gives you no speaker switch: the request schema states that provider knobs such as diarize and tag_audio_events are fixed server-side, and the documented output is words[] plus optional sentence segments.
Microsoft's claim is from its announcement post; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. This post does not say what the fixed server-side values are, because the docs do not.
What does Microsoft describe?
The post says multi-party audio, such as meetings, interviews, contact center calls and panel recordings, "comes back as structured, speaker-labeled dialogue instead of an undifferentiated block of text". It calls diarization a heavily requested capability. The announcement shows no request parameter or output field for speakers.
What does Sume expose instead?
The schema is closed: audio_url, optional segmentation, language_code, duration_seconds, metadata and the usual job mode fields. Nothing in it turns speaker separation on or off. The documented per-word entry is { word, start, end }, with no speaker property.
| Need | MAI-Transcribe-2 post | Sume `sume/stt-1.0` |
|---|---|---|
| Speaker-labeled segments | Yes | Not documented |
| Caller diarize switch | Not shown | None; knobs fixed server-side |
| Per-word timing | Yes | words[] with word, start, end |
| Sentence segments | Not stated | Optional segmentation.mode: "sentence" |
Can I still separate speakers on Sume?
Not from the STT result. If you already know who speaks when, you can cut the audio into ranges per speaker and transcribe each piece, then label the output yourself. That costs one request per piece, and the duration_seconds hint tops out at 10 minutes. Detaching audio first gets you the right input shape.
When is a different service the right call?
If your product needs speaker-attributed transcripts out of the box, such as contact center review, pick a service that documents that output. If you need accurate word times for captions and cuts, the Sume words array covers that. The related Gemini diarization post draws the same line for another vendor.
Sources
Related posts
More in Models
- Midjourney edit model image references: 4 vs Sume's ranges
Midjourney's V8.2 edit model takes up to 4 image references. On Sume the ceiling is per model: GPT Image 2.5 takes up to 16. Read the descriptor first.
- MiniMax-H3 Fun ControlNet Union 2.0: 8 conditions vs Sume references
MiniMax-H3-Fun-Controlnet-Union-2.0 adds Scribble, Layout and Gray to five older conditions. Sume's minimax-h3 ids take image, video and audio references.
- MiniMax H3 license for an EU company: contact MiniMax first
MiniMax H3 license Section II invites people in the EU, UK, Korea and USA to contact MiniMax about a license. What it promises and omits.
- MiniMax H3 license: EU, UK, Korea and US are Excluded Territories
The MiniMax H3 Community License defines the EU, UK, Republic of Korea and USA as Excluded Territories and says use there is not authorized. What it quotes.
Written by Sume