MAI-Transcribe-2-Streaming Realtime API: events vs Sume job URLs
Microsoft's streaming transcriber uses a WebSocket with delta, intermediate and commit events. Sume STT takes a file URL and returns a job. A side-by-side.

MAI-Transcribe-2-Streaming is used over a WebSocket at wss://{resource}.services.ai.azure.com/mai/v1/realtime, with an OpenAI-Realtime-like protocol: you send base64 PCM16 chunks with input_audio_buffer.append, you receive delta and intermediate events while the person talks, and you send input_audio_buffer.commit to get a completed transcript. Sume STT works the other way round: you give it a public HTTPS audio URL and read a finished transcript from a job.
That is the real difference, and it decides which of the two you can use for which task. Microsoft marks the streaming model as a public preview without a service-level agreement on both Learn pages I read on 2026-10-03.
What are the protocol facts?
These are straight from Microsoft's Learn page for the Realtime API.
| Item | What Learn says |
|---|---|
| Transport | WebSocket, OpenAI Realtime API-like protocol |
| Audio | Raw signed little-endian PCM16, mono, 16000 or 24000 samples per second, no WAV header |
| Chunk size | 10-20 ms chunks to minimise latency |
| Session length | Up to one hour per session |
| Turn detection | Only null: no server-side speech detection and no automatic commit; the client sends commit |
| Events | delta (newly final text), intermediate (MAI-specific provisional suffix), committed, completed |
| Language | Optional code; null means automatic detection across 60 languages |
What does Sume STT do with the same audio?
POST /v1/stt-1.0/transcribe takes audio_url (public HTTPS), an optional language_code (omit it for auto-detect), and an optional duration_seconds from 1 to 600. Without duration_seconds Sume reserves one minute, so send it for anything longer. A completed job returns text, language fields when available, and words[] with word, start and end in seconds from the start of the audio.
There are no partial results and no commit. mode: sync blocks for up to 30 seconds and otherwise returns polling URLs; mode: webhook delivers one terminal callback. The 30 seconds bound the HTTP wait, not the job.
import os, time, requests
BASE, H = "https://api.sume.com", {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = dict(H, **({"Idempotency-Key": key} if key else {}))
d = requests.post(BASE + path, headers=h, json=body, timeout=60)
d.raise_for_status()
d = d.json()["data"]
while not d["terminal"]:
time.sleep(d.get("next_poll_after_seconds") or 2)
d = requests.get(d["status_url"], headers=H, timeout=30).json()["data"]
r = requests.get(d["result_url"], headers=H, timeout=30)
r.raise_for_status() # failed or canceled jobs answer 409 here
return r.json()["data"]["result"]
result = run("/v1/stt-1.0/transcribe", {
"audio_url": os.environ["AUDIO_URL"],
"duration_seconds": 120,
"mode": "async",
})
print(result["text"])
print(result["words"][:3])
Which events have a Sume equivalent?
Almost none, and that is fine for the right workload. A streaming session is for a conversation in progress; a job is for a recording that already exists.
deltaandintermediatehave no equivalent: Sume never sends text before the job completes.commithas no equivalent: a job covers the whole file you gave it.completedmaps to the job result:textpluswords[].- Session length maps to a 600-second cap per STT job; Transcribe a three-hour recording shows how to split a longer file.
- Your PCM16 chunks do not map either: Sume wants a URL it can fetch, not bytes on a socket.
What would a bridge between the two look like?
Some teams want both: a live transcript while a call runs, and a clean file transcript afterwards. A workable layout is to keep the streaming socket for the live view only, record the call audio on your side to a file, and send that file to a Sume STT job when the call ends. The streaming text is then a convenience, and the file transcript, with its words[] timings, is the record you search, caption or check.
The two transcripts will not match word for word, because they come from different models and different passes. Decide up front which one is canonical. If an agent compliance check or a caption burn-in reads the transcript, point it at the job result, since its word times are tied to the file you can replay.
Remember the cap. A call longer than 600 seconds needs splitting before the STT job, and a call longer than an hour needs a second streaming session on Microsoft's side. Both limits are in the tables above, and the three-hour recording post shows the split.
How do you choose?
If a person is talking now and your product reacts as they speak, you need a streaming service, and Sume is not one; see real-time speech-to-text API for what a file API can approximate with short chunks. If the audio is a finished file, such as a call recording, a podcast or the audio track of a video, a Sume job gives you word timings you can burn into captions, check against a script, or search, and it needs no socket code or Azure resource.
Sources
Related posts
More in Developers
- Migrate Video 1.0 to sume/auto on POST /v1/videos, field by field
Video 1.0 is retiring soon. Which fields move, which are ignored or rejected, and the request to send on POST /v1/videos with sume/auto.
- Mix Seedance, Kling and Omni clips in one video: shared aspect ratio
16:9 and 9:16 are the aspect ratios Seedance 2.5, Kling 3 and Gemini Omni Flash 1.1 all list. Set the Timeline output to match and plan before render.
- Perplexity Decisions API as a publish gate for Sume output
Check a finished Sume Format image with Perplexity's Decisions API before it ships: base64 data URL, one yes/no question, a threshold, and a human-review lane.
- Pocket TTS API: run it yourself or call a hosted TTS API
Kyutai's Pocket TTS installs with pip and serves from localhost. If you want a hosted API with job URLs instead, here is the Sume request and what changes.
Written by Sume