MAI-Voice 24 kHz 160 kbps mp3 header vs Sume output_format settings
Microsoft's REST example asks for audio-24khz-160kbitrate-mono-mp3. Sume has no 160 kbps option: mp3 bit rates are 32, 64, 96, 128 or 192 kbps at up to 48 kHz.

The MAI voices REST example sets the header X-Microsoft-OutputFormat: audio-24khz-160kbitrate-mono-mp3, one string that bundles rate, bit rate, channels and container. Sume uses an output_format object instead, and its mp3 bit rates are 32, 64, 96, 128 or 192 kbps, so 160 kbps has no exact match: choose 128 or 192 kbps, and 24000 Hz is a valid sample rate.
The two ways of saying it
On the Microsoft side, the format is a single header value in the example on the voices page. On Sume, output_format takes container (mp3, wav or raw), sample_rate (8000, 16000, 22050, 24000, 44100 or 48000), bit_rate for mp3, and encoding for PCM (pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw). The default is mp3 at 44100 Hz and 128 kbps.
| Part of the MAI header | Meaning | Sume output_format |
|---|---|---|
| audio-24khz | 24000 Hz | sample_rate: 24000 |
| 160kbitrate | 160 kbps | No exact value; bit_rate 128000 or 192000 |
| mono | One channel | No channel field in this object |
| mp3 | Container | container: mp3 |
A request body that lands closest
Here is the nearest body: 24 kHz mp3 at 192 kbps, which is slightly above the MAI example and keeps quality margin. Choose 128000 when file size matters more.
import json
body = {
"transcript": "Format check.",
"language": "en",
"voice": {"id": "REPLACE_WITH_A_VOICE_ID"},
"output_format": {
"container": "mp3",
"sample_rate": 24000,
"bit_rate": 192000,
},
}
print(json.dumps(body, indent=2))
Limits
The Sume schema has no channel field in output_format, so do not assume you can ask for stereo. Treat the voice output as speech and check the result in your player. The mp3 container also adds encoder priming when you later join pieces; the timeline audio docs describe that padding for mp3 output, which is why wav is the safer choice for editing.
The MAI models are in public preview, so the header value may change; the Sume values come from the current request schema.
Choosing settings in practice
Match the setting to the destination, not to the largest number. A narrow, speech-only track gains little from a very high bit rate, and a file that will be edited again should avoid lossy re-encoding where it can. Both services expose these choices, but only the Sume object lists exactly which values are allowed, so read the enum before you pick.
- Voice-over for video: the default 44100 Hz at 128 kbps is plenty, and
wavwithpcm_s16lesuits editing. - Telephony:
pcm_mulaworpcm_alawat 8000 Hz. - Segment slicing: use
wavorraw, sinceemit_audioslices need them and mp3 returns timings only. - Smaller files: lower
bit_rate, or drop the sample rate to 24000 if the destination allows it. - Two-step pipelines: keep the first render lossless and encode to mp3 once at the end, since each mp3 pass adds loss and priming padding.
- Check the actual result: the job result reports the audio, so confirm the format before you hand it to a downstream tool.
Sources
Related posts
More in Developers
- MAI voice ids end in a model name; Sume voice ids do not
A MAI id like en-US-Harper:MAI-Voice-2.1-Flash is not a Sume voice id. Sume takes a UUID or voi_ plus 32 hex and returns 400 invalid_voice_id for anything else.
- Make a Talking Photo Speak Spanish: TTS Language, Then Fabric
Two steps on Sume: generate Spanish speech with the language set, then animate a still with Fabric. Costs for a 30-second clip and what Avatar Video can't do.
- Map a bulk queue item index back to a SKU: keep your own ledger
Queue items return index, status and run_id, not your SKU. Save a SKU-by-index table at submit time and join it to the queue receipt when you poll.
- Marketing API v24.0 ends Oct 6, 2026: pin the version in your uploader
Meta lists Marketing API v24.0 as available until October 6, 2026 and v25.0 as latest. Make the version a setting, and keep Sume render jobs separate.
Written by Sume