MAI-Transcribe-2-Streaming partials vs Sume STT terminal webhook

MAI-Transcribe-2-Streaming sends intermediate and final results as audio streams in. Sume STT sends one signed terminal callback. How to design for each.

4 min readSume
All posts

MAI-Transcribe-2-Streaming returns two kinds of result while audio is still arriving: intermediate results that update the current transcription, and final results that confirm each segment. Sume STT returns nothing until the job ends, then delivers one terminal result by poll or by a single signed webhook. Pick the stream for live captions, and the job for recordings.

What Microsoft describes

Microsoft's Learn page, last updated 2026-10-01, calls the model a low-latency speech-to-text model for real-time transcription, and marks it public preview without a service-level agreement. You choose between an OpenAI Realtime-compatible WebSocket integration and the Azure Speech SDK. The launch post lists 60 languages with continuous language detection and an introductory $0.54 per audio hour through the end of 2026 (read 2026-10-07).

MAI-Transcribe-2-Streaming facts from Microsoft pages (read 2026-10-07)
ItemDetail
StatusPublic preview, no SLA
ResultsIntermediate and final
IntegrationsRealtime API (WebSocket) or Azure Speech SDK
Languages60, with automatic continuous detection
Introductory price$0.54 per audio hour through end of 2026

What Sume STT delivers

POST /v1/stt-1.0/transcribe takes a public HTTPS audio_url of up to 600 seconds as a declared duration. With mode: webhook and a webhook_url, Sume delivers only terminal events: job.completed, job.failed and job.canceled. There are no progress or partial callbacks, so Sume's docs advise keeping status_url polling as a backup.

The result has text, words[] with start and end seconds, and, if you asked for segmentation, gapless sentence segments.

Design differences to plan for

  • State: a stream client must replace the last partial with the next; a Sume client stores one final record.
  • Edits: a partial can change words you already showed. A Sume transcript never changes after the job completes.
  • Failure: a dropped socket loses live text. A failed Sume job gives a job.failed event and a status you can re-read.
  • Verification: a webhook carries x-sume-webhook-signature with sume-v1=<hex> over <timestamp>.<raw_body>. Reject callbacks outside a replay window of about five minutes.

Where each fits

A webinar with captions on screen needs the stream. A podcast, a call recording or a clip for dubbing is already finished, and the job is simpler. Sume does not offer streaming STT, so if live is a hard requirement, Sume is the wrong layer; you can still caption the recording afterward.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: call-0001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/example/clip.m4a",
    "duration_seconds": 480,
    "mode": "webhook",
    "webhook_url": "https://example.com/webhooks/sume"
  }'

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume