MAI-Transcribe-2-Streaming partials vs Sume STT terminal webhook
MAI-Transcribe-2-Streaming sends intermediate and final results as audio streams in. Sume STT sends one signed terminal callback. How to design for each.

MAI-Transcribe-2-Streaming returns two kinds of result while audio is still arriving: intermediate results that update the current transcription, and final results that confirm each segment. Sume STT returns nothing until the job ends, then delivers one terminal result by poll or by a single signed webhook. Pick the stream for live captions, and the job for recordings.
What Microsoft describes
Microsoft's Learn page, last updated 2026-10-01, calls the model a low-latency speech-to-text model for real-time transcription, and marks it public preview without a service-level agreement. You choose between an OpenAI Realtime-compatible WebSocket integration and the Azure Speech SDK. The launch post lists 60 languages with continuous language detection and an introductory $0.54 per audio hour through the end of 2026 (read 2026-10-07).
| Item | Detail |
|---|---|
| Status | Public preview, no SLA |
| Results | Intermediate and final |
| Integrations | Realtime API (WebSocket) or Azure Speech SDK |
| Languages | 60, with automatic continuous detection |
| Introductory price | $0.54 per audio hour through end of 2026 |
What Sume STT delivers
POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url of up to 600 seconds as a declared duration. With mode: webhook and a webhook_url, Sume delivers only terminal events: job.completed, job.failed and job.canceled. There are no progress or partial callbacks, so Sume's docs advise keeping status_url polling as a backup.
The result has text, words[] with start and end seconds, and, if you asked for segmentation, gapless sentence segments.
Design differences to plan for
- State: a stream client must replace the last partial with the next; a Sume client stores one final record.
- Edits: a partial can change words you already showed. A Sume transcript never changes after the job completes.
- Failure: a dropped socket loses live text. A failed Sume job gives a
job.failedevent and a status you can re-read. - Verification: a webhook carries
x-sume-webhook-signaturewithsume-v1=<hex>over<timestamp>.<raw_body>. Reject callbacks outside a replay window of about five minutes.
Where each fits
A webinar with captions on screen needs the stream. A podcast, a call recording or a clip for dubbing is already finished, and the job is simpler. Sume does not offer streaming STT, so if live is a hard requirement, Sume is the wrong layer; you can still caption the recording afterward.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: call-0001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/example/clip.m4a",
"duration_seconds": 480,
"mode": "webhook",
"webhook_url": "https://example.com/webhooks/sume"
}'Sources
Related posts
More in Comparisons
- MAI custom voice SSML (ttsembedding) vs Sume voice ids: what is gated
Microsoft's MAI-Voice-2.1 custom voice sits behind Limited Access Review and uses a speakerProfileId in SSML. Sume picks voices by id, with no clone widget.
- Edit models Morphic names versus what the Sume catalog lists
Morphic runs image edits on Seedream 5.0 Pro, Nano Banana 2 and GPT Image 2.5 until Ideogram 4.5 arrives. Which of these Sume lists, and where to check.
- AI music cost per track: Sume Music vs Lyria 3.5 and ACE Step
Sume Music is a flat $0.125 per generation ($12.50 per 100). Google lists Lyria 3.5 at $0.08 per song ($8.00); fal lists ACE Step at $0.0002 per second.
- Nano Banana 2.1 costs the same as Lite at 1K ($0.0336): what differs
At 1K, Nano Banana 2.1 and Nano Banana 2 Lite both list at $0.0336 per image. The gap is in input and thinking tokens, plus what Lite cannot do.
Written by Sume