MAI-Transcribe-2-Streaming Realtime API: events vs Sume job URLs

Microsoft's streaming transcriber uses a WebSocket with delta, intermediate and commit events. Sume STT takes a file URL and returns a job. A side-by-side.

5 min readSume
All posts

MAI-Transcribe-2-Streaming is used over a WebSocket at wss://{resource}.services.ai.azure.com/mai/v1/realtime, with an OpenAI-Realtime-like protocol: you send base64 PCM16 chunks with input_audio_buffer.append, you receive delta and intermediate events while the person talks, and you send input_audio_buffer.commit to get a completed transcript. Sume STT works the other way round: you give it a public HTTPS audio URL and read a finished transcript from a job.

That is the real difference, and it decides which of the two you can use for which task. Microsoft marks the streaming model as a public preview without a service-level agreement on both Learn pages I read on 2026-10-03.

What are the protocol facts?

These are straight from Microsoft's Learn page for the Realtime API.

MAI-Transcribe-2-Streaming Realtime API facts from Microsoft Learn, read 2026-10-03.
ItemWhat Learn says
TransportWebSocket, OpenAI Realtime API-like protocol
AudioRaw signed little-endian PCM16, mono, 16000 or 24000 samples per second, no WAV header
Chunk size10-20 ms chunks to minimise latency
Session lengthUp to one hour per session
Turn detectionOnly null: no server-side speech detection and no automatic commit; the client sends commit
Eventsdelta (newly final text), intermediate (MAI-specific provisional suffix), committed, completed
LanguageOptional code; null means automatic detection across 60 languages

What does Sume STT do with the same audio?

POST /v1/stt-1.0/transcribe takes audio_url (public HTTPS), an optional language_code (omit it for auto-detect), and an optional duration_seconds from 1 to 600. Without duration_seconds Sume reserves one minute, so send it for anything longer. A completed job returns text, language fields when available, and words[] with word, start and end in seconds from the start of the audio.

There are no partial results and no commit. mode: sync blocks for up to 30 seconds and otherwise returns polling URLs; mode: webhook delivers one terminal callback. The 30 seconds bound the HTTP wait, not the job.

import os, time, requests
BASE, H = "https://api.sume.com", {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
    h = dict(H, **({"Idempotency-Key": key} if key else {}))
    d = requests.post(BASE + path, headers=h, json=body, timeout=60)
    d.raise_for_status()
    d = d.json()["data"]
    while not d["terminal"]:
        time.sleep(d.get("next_poll_after_seconds") or 2)
        d = requests.get(d["status_url"], headers=H, timeout=30).json()["data"]
    r = requests.get(d["result_url"], headers=H, timeout=30)
    r.raise_for_status()  # failed or canceled jobs answer 409 here
    return r.json()["data"]["result"]

result = run("/v1/stt-1.0/transcribe", {
    "audio_url": os.environ["AUDIO_URL"],
    "duration_seconds": 120,
    "mode": "async",
})
print(result["text"])
print(result["words"][:3])

Which events have a Sume equivalent?

Almost none, and that is fine for the right workload. A streaming session is for a conversation in progress; a job is for a recording that already exists.

  • delta and intermediate have no equivalent: Sume never sends text before the job completes.
  • commit has no equivalent: a job covers the whole file you gave it.
  • completed maps to the job result: text plus words[].
  • Session length maps to a 600-second cap per STT job; Transcribe a three-hour recording shows how to split a longer file.
  • Your PCM16 chunks do not map either: Sume wants a URL it can fetch, not bytes on a socket.

What would a bridge between the two look like?

Some teams want both: a live transcript while a call runs, and a clean file transcript afterwards. A workable layout is to keep the streaming socket for the live view only, record the call audio on your side to a file, and send that file to a Sume STT job when the call ends. The streaming text is then a convenience, and the file transcript, with its words[] timings, is the record you search, caption or check.

The two transcripts will not match word for word, because they come from different models and different passes. Decide up front which one is canonical. If an agent compliance check or a caption burn-in reads the transcript, point it at the job result, since its word times are tied to the file you can replay.

Remember the cap. A call longer than 600 seconds needs splitting before the STT job, and a call longer than an hour needs a second streaming session on Microsoft's side. Both limits are in the tables above, and the three-hour recording post shows the split.

How do you choose?

If a person is talking now and your product reacts as they speak, you need a streaming service, and Sume is not one; see real-time speech-to-text API for what a file API can approximate with short chunks. If the audio is a finished file, such as a call recording, a podcast or the audio track of a video, a Sume job gives you word timings you can burn into captions, check against a script, or search, and it needs no socket code or Azure resource.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume