Send a WAV file to MAI-Transcribe-2-Streaming: PCM16, no WAV header

MAI streaming wants raw PCM16 mono at 16 or 24 kHz as base64 in 10-20 ms chunks, no header. Python's wave module strips it; Sume's STT shape is 16 kHz mono.

4 min readSume
All posts

To stream a WAV file into MAI-Transcribe-2-Streaming, read its frames with Python's wave module and send them as base64 of raw signed 16-bit little-endian mono PCM at 16000 or 24000 Hz, with no WAV header. Microsoft suggests chunks of 10 to 20 ms. The same 16 kHz mono shape is what Sume's own audio-detach docs call the STT shape.

What the endpoint expects

The realtime page lists the input format and the endpoint wss://{resource}.services.ai.azure.com/mai/v1/realtime?intent=transcription. A header left in the first chunk would be read as audio, which is why the page is explicit about it. You also send input_audio_buffer.commit yourself, because server turn detection is off in the preview.

MAI streaming audio input and the Sume equivalent (Microsoft page read 2026-10-05; Sume docs)
PropertyMAI-Transcribe-2-StreamingSume STT 1.0
ContainerNone; raw PCM16 in base64Public HTTPS audio file URL
ChannelsMono16 kHz mono is the documented STT shape
Sample rate16000 or 24000 HzAudio detach can output 16000 Hz
Chunk size10-20 ms suggestedNot applicable

Chunk a WAV into 20 ms slices

This script checks the file is mono, 16-bit and at an accepted rate, then yields base64 chunks. It makes no network call, so you can test it offline.

import base64, sys, wave

OK_RATES = (16000, 24000)

def chunks(path, ms=20):
    with wave.open(path, "rb") as w:
        if w.getnchannels() != 1 or w.getsampwidth() != 2:
            raise SystemExit("need mono 16-bit PCM")
        rate = w.getframerate()
        if rate not in OK_RATES:
            raise SystemExit(f"rate {rate} not accepted")
        step = rate * ms // 1000
        while True:
            frames = w.readframes(step)
            if not frames:
                return
            yield base64.b64encode(frames).decode()

if __name__ == "__main__":
    n = sum(1 for _ in chunks(sys.argv[1]))
    print(n, "chunks of 20 ms")

Mistakes this catches

Most failed first attempts at a raw-audio socket come from a short list of causes. A header sent inside the first chunk reads as a burst of noise. Stereo sent as mono doubles the playback speed of what the model hears. A sample rate that is neither 16000 nor 24000 is outside the documented input. Big-endian samples are wrong too, since the page specifies little-endian.

The script above checks channel count, sample width and rate before it reads a single frame, and it only ever returns frames, never header bytes.

  • Chunk length is a trade-off: smaller chunks mean lower latency but more events to handle, so start at 20 ms and measure.
  • Send input_audio_buffer.commit after the last chunk, since there is no server VAD to do it for you.
  • Do not change language or format mid-session; the session settings lock after the first append.

Getting a file into that shape

If your source is a video, audio detach can output a 16 kHz mono track; the docs describe sample_rate: 16000 with channels: "mono" as the STT shape. That helps for Sume STT as well as for this script.

Limits: a WAV with a 44.1 kHz rate or stereo will exit the script rather than be silently converted, on purpose. Resample first. Also remember the preview has no SLA and a one-hour session cap, so treat this as a test harness before a production path.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume