MAI-Transcribe-2 word timestamps vs Sume STT words[] shape

MAI-Transcribe-2 gives every word a start and end time. Sume STT always returns words[] as { word, start, end } in seconds, with no flag to turn it on.

4 min readSume
All posts

Microsoft says MAI-Transcribe-2 gives every word its own start and end time. Sume STT does the same by default: the result carries words[], each entry { word, start, end } in seconds from the start of the audio, and there is no flag to enable it.

Microsoft's wording is from its announcement post; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. For the full tour of the words array see speech-to-text word timestamps.

What does Microsoft promise for timestamps?

The post lists word-level timestamps as one of two structural capabilities: every word carries a start and an end time. It names click-to-play navigation, frame-accurate captioning and subtitling, redaction of spoken segments and search that jumps to a term as uses. It does not print a response schema, so the exact field names and units are not in the post.

What does a Sume word entry contain?

Each entry in words[] has word, start and end, in seconds from the audio start. The request schema says word timings are always returned, so there is no flag to enable them. Time values are plain numbers you can feed straight into a caption or cut list.

Word-timing facts, Microsoft post vs Sume STT schema, read 2026-10-01.
QuestionMAI-Transcribe-2 postSume `sume/stt-1.0`
Per-word start and endYes, "every word carries its own start and end time"Yes, { word, start, end }
UnitNot statedSeconds from audio start
Needs a flagNot statedNo, always returned
Sentence groupingNot statedOptional segmentation

How do I get sentence-sized captions instead?

Add segmentation: { "mode": "sentence" }. Sume then derives gapless sentence segments from the word timings, grouping on terminal punctuation and splitting unpunctuated runs on silence. If the provider returned no timed words the request fails closed with a typed error rather than guessing.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-demo-001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
    "duration_seconds": 120,
    "segmentation": { "mode": "sentence" }
  }'

Is the output accurate to the frame?

Neither source here measures it, and this post makes no accuracy claim. Check your own clips against a few known words before you build frame-accurate captions on top of either service.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume