Speaker diarization API: confidence scores, and what Sume STT returns
AssemblyAI added speaker_confidence (0 to 1) per word and utterance. Sume STT returns word start and end times only, with no speaker or confidence data.

A speaker diarization API labels which speaker said each part of a recording. AssemblyAI's changelog says that, when you set speaker_options.include_speaker_confidence to true, each word and each utterance gets a speaker_confidence value between 0 and 1. Sume STT returns word start and end times and optional sentence segments, with no speaker label and no confidence value.
AssemblyAI facts are from its changelog, read 2026-09-30. Sume facts are from the OpenAPI document for sume/stt-1.0.
What is a speaker confidence score for?
The changelog says it shows how confident the model is in a given speaker assignment. In practice you would use it to route low-score segments to a human reviewer, or to skip them when you build per-speaker clips. It only makes sense where a model assigns speakers in the first place.
What fields does a Sume STT result have?
| Field | AssemblyAI diarization | Sume STT 1.0 |
|---|---|---|
| Speaker label | Per utterance | None |
| Confidence | speaker_confidence, 0 to 1, on request | None |
| Word times | Yes | words[] as { word, start, end } in seconds |
| Sentences | Utterances | Optional, from segmentation: { "mode": "sentence" } |
Can I ask Sume STT to diarize?
No. The request body is closed, and the MCP tool description says diarize and tag_audio_events are fixed server-side, so sending them is rejected. A flag like AssemblyAI's would be rejected the same way.
What can I build from word timings alone?
Captions, paper edits and sentence-level cuts do not need speakers; see word timestamps. For a multi-person interview, record each person on their own file, transcribe each as a separate job and interleave by start, as in how to transcribe an interview. Speaker identity then comes from your recording setup, and there is no confidence score to read.
Sources
Related posts
More in Developers
- Why Sume Agent Completions rejects assistant messages
An assistant turn in messages[] returns 400 invalid_request on Sume Agent Completions. Only system and user turns work; each call runs in a fresh thread.
- Keep the source's channels and sample rate when extracting audio
Adobe lists Match Source audio channels and sample rate in Premiere 26.5. Sume audio-detach keeps both by default: channels source, sample_rate omitted, wav.
- Extract audio from a long video: the 900 s output cap and range
Audio detach takes sources up to 1800 s but outputs at most 900 s. For a longer track pass range and detach in two calls. Errors included.
- Bluesky video upload limit: 300 MB MP4, and fitting a clip
Bluesky raised its video upload limit from 100 MB to 300 MB in June 2026, MP4 only. Fit a clip under it by trimming length and conforming size with Sume.
Written by Sume