MAI-Transcribe-2 word timestamps vs Sume STT words[] shape
MAI-Transcribe-2 gives every word a start and end time. Sume STT always returns words[] as { word, start, end } in seconds, with no flag to turn it on.

Microsoft says MAI-Transcribe-2 gives every word its own start and end time. Sume STT does the same by default: the result carries words[], each entry { word, start, end } in seconds from the start of the audio, and there is no flag to enable it.
Microsoft's wording is from its announcement post; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. For the full tour of the words array see speech-to-text word timestamps.
What does Microsoft promise for timestamps?
The post lists word-level timestamps as one of two structural capabilities: every word carries a start and an end time. It names click-to-play navigation, frame-accurate captioning and subtitling, redaction of spoken segments and search that jumps to a term as uses. It does not print a response schema, so the exact field names and units are not in the post.
What does a Sume word entry contain?
Each entry in words[] has word, start and end, in seconds from the audio start. The request schema says word timings are always returned, so there is no flag to enable them. Time values are plain numbers you can feed straight into a caption or cut list.
| Question | MAI-Transcribe-2 post | Sume `sume/stt-1.0` |
|---|---|---|
| Per-word start and end | Yes, "every word carries its own start and end time" | Yes, { word, start, end } |
| Unit | Not stated | Seconds from audio start |
| Needs a flag | Not stated | No, always returned |
| Sentence grouping | Not stated | Optional segmentation |
How do I get sentence-sized captions instead?
Add segmentation: { "mode": "sentence" }. Sume then derives gapless sentence segments from the word timings, grouping on terminal punctuation and splitting unpunctuated runs on silence. If the provider returned no timed words the request fails closed with a typed error rather than guessing.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-demo-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"duration_seconds": 120,
"segmentation": { "mode": "sentence" }
}'Is the output accurate to the frame?
Neither source here measures it, and this post makes no accuracy claim. Check your own clips against a few known words before you build frame-accurate captions on top of either service.
Sources
Related posts
More in Developers
- MAI-Voice-2.1: 23 languages and voice matching; Sume takes a voice id
Microsoft says MAI-Voice-2.1 covers 23 languages and matches a voice from a short clip. Sume TTS takes a ready voice id or avatar, not reference audio.
- MAI-Voice-2.1-Flash lists 45 ms; Sume TTS is an async job
Microsoft lists MAI-Voice-2.1-Flash at about 45 ms and MAI-Voice-2.1 at about 550 ms. Sume TTS 1.0 is non-streaming: a job with a poll URL or webhook.
- Make AI Agent fallback connection retries once: key Sume calls
Make now retries an AI Agent run once on a fallback connection. Keep the Sume Idempotency-Key out of the model's hands so a retried run cannot bill twice.
- Make a voice louder than the music: gain_db and duck_db ranges
In Timeline 1.0, raise the voice with audio.gain_db (-60 to 12) and lower the music bed with soundtrack.duck_db (0 to 20). Ranges and refusal codes.
Written by Sume