MAI-Transcribe-2 timestamps option vs Sume STT: words[] are always on

On Microsoft's MAI-Transcribe you ask for word timestamps in modelOptions. Sume STT returns words[] with start and end on every job: what that changes for you.

4 min readSume
All posts

The difference

On Microsoft's MAI-Transcribe, word timestamps are an option you request in the model options. On Sume STT you do not request them: every completed job returns words[] as {word, start, end}, in seconds from the start of the audio, and there is no flag to turn them on or off.

The practical effect is that a caption or search pipeline on Sume never has a failure where the timings were forgotten in the request. The cost is that you cannot switch them off to get a smaller response.

Side by side

Microsoft's facts below come from its MAI-Transcribe page. Sume's come from its API reference.

Word timing and related options (read 2026-10-04)
ItemMAI-Transcribe (Microsoft Learn)Sume STT 1.0
Word timestampsRequested with modelOptions.timestampsAlways returned in words[]
DiarizationOption on the requestFixed on the server; not a request field
Keyword biasingphraseList optionNo vocabulary field
Transcribe styleVerbatim or cleanNot selectable
StatusPublic preview, no SLAJob API with status and result URLs

Using words[] on Sume

Submit POST /v1/stt-1.0/transcribe with an audio_url, an optional language_code and, for the right reservation, duration_seconds between 1 and 600. The result has text and words[]. If you want sentences, add segmentation: {mode: "sentence"} and Sume groups the words on terminal punctuation, splitting unpunctuated runs on silence. The optional boundary_lead_ms defaults to 70 and carries a little lead past the last word of a sentence before the next starts, the same rule as TTS segments.

If segmentation cannot be built because the provider returned no timed words, the request fails with a typed error rather than returning guesses.

Which to choose

Choose by the controls you need. If you need keyword biasing for product names, a clean or verbatim style, or labelled speakers, Microsoft's service lists those, and the page also says it is public preview. If you want a stable, small request that always returns timings, and you accept fewer knobs, Sume STT is simpler. Price is $0.01 per audio minute on Sume.

  • Run the same ten minutes through both before you decide.
  • Check how each handles names and numbers.
  • Keep timings as floats in seconds; convert at the edge.
  • Re-check Microsoft's page: preview features change.

What to test

Timestamps look simple and are easy to get subtly wrong. Test three things on both services with the same file. First, check that the first word starts close to when you hear it, and that the last word ends near the end of the speech. Second, look at how numbers and hyphenated words are split, since a caption tool that expects one token may receive two. Third, find a place with a long pause and see whether the word before it ends where the sound ends or stretches across the silence.

If you build captions or a search index on top of the times, store the raw array with the job id and generate captions from it later. That way a change in your caption rules never requires a second paid transcription. At $0.01 a minute the transcription is cheap, but the habit of not re-running it is what keeps a large archive affordable.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume