MAI-Transcribe-2-Streaming partials vs Sume STT word timings
MAI-Transcribe-2-Streaming sends partials after about 100 ms in 60 languages. Sume STT is a job that always returns word timings. Which fits what.

What is the difference between a streaming transcriber and a job-based one? Microsoft says MAI-Transcribe-2-Streaming returns its first hypotheses, called partials, in just over 100 ms and covers 60 languages. Sume STT 1.0 takes a public audio URL, runs as a job and always returns word-level timings.
The first suits a live caption or dictation box. The second suits a finished recording that you want to cut, subtitle or search.
Partials and finals
A streaming model emits provisional text that it later revises, then a final transcript. Microsoft reports accuracy rankings for both partial and final transcripts on Artificial Analysis, and an introductory price of $0.54 per hour of audio through the end of the year.
A job-based transcriber has no provisional stage. You submit a file and read one result when the job completes, with no text on screen while it runs.
| Item | MAI-Transcribe-2-Streaming | Sume STT 1.0 |
|---|---|---|
| Delivery | Streaming partials, then finals | A job; read the result by polling or webhook |
| First text | Just over 100 ms for partials | After the job completes |
| Languages | 60, per Microsoft | language_code hint, or auto-detect when omitted |
| Price | $0.54 per hour of audio, introductory, through year end | $0.01 per audio minute, which is $0.60 per hour |
| Word timings | Not described on the cited page | Always returned |
| Sentences | Not described on the cited page | Optional segmentation.mode sentence |
What Sume returns
The Sume request needs a public HTTPS audio_url. language_code is optional and omitted means auto-detect. duration_seconds, from 1 to 600, only improves the usage reservation; leave it out and Sume reserves one minute. Word timings come back on every request, with no flag, and adding segmentation also returns sentence segments with a default 70 ms lead past each sentence's last word, adjustable from 0 to 500 ms.
The hourly figure in the table is plain arithmetic: $0.01 times 60 minutes. The two prices are not like for like, since one is a streaming service in preview terms and the other is a job with a catalog rate; compare them for your workload rather than as a ranking.
Choosing by the job you have
Pick on what the user sees, not on the headline price.
- Live captions, dictation or an agent that must hear a caller as they speak: a streaming model, because a job cannot show text while the speaker talks.
- Subtitles, clips and search over a finished video: a job, because timings for every word are what you need, and the delay is invisible.
- Both: stream during the session, then run the saved recording through a job for the clean timed transcript you publish.
Waiting for a job
A Sume job can be waited on for up to 30 seconds in sync mode, or polled, or delivered by webhook, as described in Sume jobs and results. For a long recording use webhook or polling rather than a held request. The cost side of the same comparison is in the cost of one hour.
Sources
Related posts
More in Comparisons
- Make an AI avatar say exact words: Tavus echo mode vs Sume
Tavus echo mode sends text or audio straight to the avatar for playback, skipping perception and speech recognition. Sume's avatar video renders your script.
- Micro-drama lead: Avatar 1.0 or Seedance 2.5? Pick by dialogue
Pick Sume Avatar 1.0 for talking to camera and Seedance 2.5 for a lead who moves through places. Cost per minute: $14.70 against $34.67.
- Clipchamp free captions and silence removal vs a scripted cut list
Clipchamp's pricing page lists AI subtitles and silence removal in the free plan. When a scripted Sume cut list still makes sense, and what each step costs.
- Midjourney alternative for product stills by API: what to send on Sume
Sume's image catalog has no Midjourney model. For product stills it lists reference-based edit models instead; here is which id fits which job, and how to test.
Written by Sume