MAI-Transcribe-2 timestamps option vs Sume STT: words[] are always on
On Microsoft's MAI-Transcribe you ask for word timestamps in modelOptions. Sume STT returns words[] with start and end on every job: what that changes for you.

The difference
On Microsoft's MAI-Transcribe, word timestamps are an option you request in the model options. On Sume STT you do not request them: every completed job returns words[] as {word, start, end}, in seconds from the start of the audio, and there is no flag to turn them on or off.
The practical effect is that a caption or search pipeline on Sume never has a failure where the timings were forgotten in the request. The cost is that you cannot switch them off to get a smaller response.
Side by side
Microsoft's facts below come from its MAI-Transcribe page. Sume's come from its API reference.
| Item | MAI-Transcribe (Microsoft Learn) | Sume STT 1.0 |
|---|---|---|
| Word timestamps | Requested with modelOptions.timestamps | Always returned in words[] |
| Diarization | Option on the request | Fixed on the server; not a request field |
| Keyword biasing | phraseList option | No vocabulary field |
| Transcribe style | Verbatim or clean | Not selectable |
| Status | Public preview, no SLA | Job API with status and result URLs |
Using words[] on Sume
Submit POST /v1/stt-1.0/transcribe with an audio_url, an optional language_code and, for the right reservation, duration_seconds between 1 and 600. The result has text and words[]. If you want sentences, add segmentation: {mode: "sentence"} and Sume groups the words on terminal punctuation, splitting unpunctuated runs on silence. The optional boundary_lead_ms defaults to 70 and carries a little lead past the last word of a sentence before the next starts, the same rule as TTS segments.
If segmentation cannot be built because the provider returned no timed words, the request fails with a typed error rather than returning guesses.
Which to choose
Choose by the controls you need. If you need keyword biasing for product names, a clean or verbatim style, or labelled speakers, Microsoft's service lists those, and the page also says it is public preview. If you want a stable, small request that always returns timings, and you accept fewer knobs, Sume STT is simpler. Price is $0.01 per audio minute on Sume.
- Run the same ten minutes through both before you decide.
- Check how each handles names and numbers.
- Keep timings as floats in seconds; convert at the edge.
- Re-check Microsoft's page: preview features change.
What to test
Timestamps look simple and are easy to get subtly wrong. Test three things on both services with the same file. First, check that the first word starts close to when you hear it, and that the last word ends near the end of the speech. Second, look at how numbers and hyphenated words are split, since a caption tool that expects one token may receive two. Third, find a place with a long pause and see whether the word before it ends where the sound ends or stretches across the silence.
If you build captions or a search index on top of the times, store the raw array with the job id and generate captions from it later. That way a change in your caption rules never requires a second paid transcription. At $0.01 a minute the transcription is cheap, but the habit of not re-running it is what keeps a large archive affordable.
Sources
Related posts
More in Comparisons
- MAI-Voice-2.1 lists Korean, Thai, Vietnamese: Sume's language field
Microsoft's MAI-Voice-2.1 page lists 23 languages including Korean, Thai and Vietnamese. What to check before you plan a non-English voiceover on Sume TTS.
- MAI-Voice-2.1 or Flash for explainer narration: latency barely matters
Microsoft lists about 550 ms vs 45 ms model inference. For narration rendered ahead of time, pick on quality and price, then see how Sume jobs return audio.
- MakeUGC API Starter $99 for 2,000 credits vs Sume avatar seconds
MakeUGC API Starter is $99 a month with 2,000 credits. The same $99 on Sume buys 13 Plus 30-second avatar clips at $7.35 each, or 17 on Standard at $5.52.
- Meshy 7.1 Ultra 4K: when a 3D model is overkill
Meshy 7.1 added Ultra 4K meshes of up to 80 million triangles. Sume has no image-to-3D model; to show a product turning, a short video does it.
Written by Sume