MAI-Transcribe-2-Streaming price: $0.54 an hour vs batch vs Sume
Microsoft lists streaming transcription at $0.54 an hour through year end; batch MAI-Transcribe-2 is reported at $0.10. Sume STT is $0.01 a minute. The math.

MAI-Transcribe-2-Streaming costs $0.54 per hour of audio through the end of 2026, according to Microsoft's announcement. That is $0.009 a minute. Sume STT bills $0.01 per audio minute, which is $0.60 an hour. So on list price streaming MAI is about 10% cheaper per hour than Sume's file transcription, and it does a different job.
The batch MAI-Transcribe-2 price of $0.10 an hour comes from a news report, not from a page I could read on Microsoft's site, so it is marked "reported" below. If your audio is a recording rather than a live call, the batch price is the one to compare, and Sume is six times higher on list.
What are the three prices side by side?
Per-hour figures are computed from each per-minute or per-hour list price. Sume's $0.01 a minute is the pricing code's STT provider rate times Sume's margin, and duration_seconds sets the reservation.
| Service | Listed price | Per hour | Notes |
|---|---|---|---|
| MAI-Transcribe-2-Streaming | $0.54 per hour of audio | $0.54 | Introductory through the end of 2026; public preview |
| MAI-Transcribe-2 (batch) | $0.10 per hour (reported) | $0.10 | Reported by a news site; not verified on a Microsoft page |
| Sume STT 1.0 | $0.01 per audio minute | $0.60 | File URL in, job out; up to 600 s per job |
Why is streaming dearer than batch?
Streaming returns partial text within about a tenth of a second (Microsoft's claim of just over 100 ms to first partials) which a batch job does not promise. I could not find Microsoft explaining the price gap, so I will not guess at it; the Learn page points to the Foundry Models pricing page for current prices, and says the model is in public preview with no service-level agreement, so check that page before you commit.
How do you price a month of audio?
Multiply hours by the rate, and compare both lines for your real mix of live and recorded audio. The script below does it for a given number of hours; the rates are the ones in the table.
RATES = {
"mai_streaming": 0.54, # USD per hour, introductory
"mai_batch": 0.10, # USD per hour, reported
"sume_stt": 0.60, # USD per hour ($0.01 per minute)
}
def month(live_hours, recorded_hours):
return {
"streaming for live, batch for recorded": round(
live_hours * RATES["mai_streaming"] + recorded_hours * RATES["mai_batch"], 2),
"streaming for everything": round((live_hours + recorded_hours) * RATES["mai_streaming"], 2),
"Sume for recorded only": round(recorded_hours * RATES["sume_stt"], 2),
}
print(month(live_hours=200, recorded_hours=800))
What does the introductory price mean for planning?
Microsoft's page states the $0.54 rate as introductory and as available through the end of the year. Anything you budget for 2027 should assume it can change, and the Learn pages send you to the Foundry Models pricing page for the current figure. The preview label matters too: both Learn pages I read say the feature is in public preview, without a service-level agreement, and not recommended for production workloads.
For a Sume comparison the opposite holds. The STT rate comes from the pricing code on main and the OpenAPI document, so it is a current figure, but it can change in a release. Read GET /v1/catalog for the live price before you build a model on it.
Either way, build your cost model as a rate table you can edit, as in the snippet, not as constants spread through your code. For scale, the script's sample input of 200 live hours and 800 recorded hours prices the mixed plan at $188 a month, streaming for everything at $540, and Sume for the recorded 800 hours alone at $480. Those figures only show how the choice of service moves the total; your mix of live and recorded audio is the number to change.
When is Sume's higher rate still the right call?
When volume is large and the audio is a plain recording, run the numbers: cost to transcribe 1,000 hours of audio shows the gap at scale, and live vs file transcription cost covers the same split for OpenAI.
- You need no socket code, no Azure resource and no regional deployment, only an HTTPS URL and one key.
- You want
words[]timings you can feed straight into caption burn-in or script alignment on Sume. - Your volume is small enough that the per-hour gap is a few dollars a month.
- You want one invoice for transcription, voice, captions and video.
Sources
Related posts
More in Pricing
- Mercury Voice pricing: tokens to dollars per minute of talk
Mercury Voice lists at $0.40/$1.50 per million tokens, half off at launch, about $0.009 a minute. How to turn that into your bill, plus the speech steps.
- Balance needed for one AI generation call: published max per endpoint
Sume's catalog publishes estimated, minimum and maximum cents per endpoint. The worst-case hold for one call runs from 3 cents to $56.25 before a 402 can fire.
- Real-time AI avatar pricing: live minutes vs one render
Live avatars bill by conversation minute and concurrent stream; a rendered clip is paid once and watched by anyone. Tavus plan numbers and the break-even.
- Sume plans: Pro $40, Startup $120, Scale $400 and what they limit
What each Sume plan sets: monthly price, concurrent jobs, queue size and API write budget. Usage is billed at each model's rate, not by the plan.
Written by Sume