MAI-Voice 24 kHz 160 kbps mp3 header vs Sume output_format settings

Microsoft's REST example asks for audio-24khz-160kbitrate-mono-mp3. Sume has no 160 kbps option: mp3 bit rates are 32, 64, 96, 128 or 192 kbps at up to 48 kHz.

4 min readSume
All posts

The MAI voices REST example sets the header X-Microsoft-OutputFormat: audio-24khz-160kbitrate-mono-mp3, one string that bundles rate, bit rate, channels and container. Sume uses an output_format object instead, and its mp3 bit rates are 32, 64, 96, 128 or 192 kbps, so 160 kbps has no exact match: choose 128 or 192 kbps, and 24000 Hz is a valid sample rate.

The two ways of saying it

On the Microsoft side, the format is a single header value in the example on the voices page. On Sume, output_format takes container (mp3, wav or raw), sample_rate (8000, 16000, 22050, 24000, 44100 or 48000), bit_rate for mp3, and encoding for PCM (pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw). The default is mp3 at 44100 Hz and 128 kbps.

Translating the header (Microsoft voices page read 2026-10-05; Sume from the repo)
Part of the MAI headerMeaningSume output_format
audio-24khz24000 Hzsample_rate: 24000
160kbitrate160 kbpsNo exact value; bit_rate 128000 or 192000
monoOne channelNo channel field in this object
mp3Containercontainer: mp3

A request body that lands closest

Here is the nearest body: 24 kHz mp3 at 192 kbps, which is slightly above the MAI example and keeps quality margin. Choose 128000 when file size matters more.

import json

body = {
    "transcript": "Format check.",
    "language": "en",
    "voice": {"id": "REPLACE_WITH_A_VOICE_ID"},
    "output_format": {
        "container": "mp3",
        "sample_rate": 24000,
        "bit_rate": 192000,
    },
}
print(json.dumps(body, indent=2))

Limits

The Sume schema has no channel field in output_format, so do not assume you can ask for stereo. Treat the voice output as speech and check the result in your player. The mp3 container also adds encoder priming when you later join pieces; the timeline audio docs describe that padding for mp3 output, which is why wav is the safer choice for editing.

The MAI models are in public preview, so the header value may change; the Sume values come from the current request schema.

Choosing settings in practice

Match the setting to the destination, not to the largest number. A narrow, speech-only track gains little from a very high bit rate, and a file that will be edited again should avoid lossy re-encoding where it can. Both services expose these choices, but only the Sume object lists exactly which values are allowed, so read the enum before you pick.

  • Voice-over for video: the default 44100 Hz at 128 kbps is plenty, and wav with pcm_s16le suits editing.
  • Telephony: pcm_mulaw or pcm_alaw at 8000 Hz.
  • Segment slicing: use wav or raw, since emit_audio slices need them and mp3 returns timings only.
  • Smaller files: lower bit_rate, or drop the sample rate to 24000 if the destination allows it.
  • Two-step pipelines: keep the first render lossless and encode to mp3 once at the end, since each mp3 pass adds loss and priming padding.
  • Check the actual result: the job result reports the audio, so confirm the format before you hand it to a downstream tool.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume