MAI-Voice REST header vs Sume TTS output_format object

MAI-Voice-2.1 sets audio format in an X-Microsoft-OutputFormat header. Sume TTS uses an output_format object. Defaults and options side by side.

5 min readSume
All posts

The two APIs ask for audio format in different places. With MAI-Voice-2.1 you send an SSML body and set the format in an X-Microsoft-OutputFormat header, for example audio-24khz-160kbitrate-mono-mp3. With Sume TTS 1.0 you send a JSON body and set an output_format object with container, sample rate, bit rate and encoding.

The Microsoft details come from its Learn page. The Sume details come from the Sume OpenAPI spec.

The two requests

Microsoft's REST example posts to https://<region>.tts.speech.microsoft.com/cognitiveservices/v1 with three headers: Content-Type: application/ssml+xml, the output format, and Ocp-Apim-Subscription-Key. The voice is named inside the SSML, such as en-US-Harper:MAI-Voice-2.1-Flash, and the response is the audio bytes.

Sume's request is JSON with a plain transcript, a voice selector and optional output_format. The response is a job; the audio is read from the finished result. A Sume example is below, with the voice taken from an environment variable.

Audio format controls, MAI-Voice REST and Sume TTS 1.0 (read 2026-10-03)
ItemMAI-Voice RESTSume TTS 1.0
Where the format goesX-Microsoft-OutputFormat headeroutput_format object in the JSON body
Documented exampleaudio-24khz-160kbitrate-mono-mp3Default mp3, 44100 Hz, 128000 bit rate
ContainersNamed in the header stringmp3, wav, raw
Sample ratesNamed in the header string8000, 16000, 22050, 24000, 44100, 48000
PCM encodingsNot shown on the pagepcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw
ResultAudio bytes in the responseA job; audio in the job result

Why the difference matters in practice

If you move from one to the other, the format string does not translate one to one. A header value like audio-24khz-160kbitrate-mono-mp3 becomes a container of mp3, a sample rate of 24000 and a bit rate; Sume's bit-rate list tops out at 192000 and the MP3 option includes 128000, but 160000 is not in the list, so the nearest accepted value must be picked deliberately.

Telephony is the other trap. Sume exposes pcm_mulaw and pcm_alaw at 8000 Hz for phone-bound audio. If your destination is a phone line, set the container and encoding explicitly rather than relying on the mp3 default.

A Sume request with an explicit format

This call asks for 24 kHz 16-bit WAV, which suits a later join or a lip-sync step. Set SUME_API_KEY and SUME_AVATAR_HANDLE first; the avatar must have a ready voice.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tts-format-demo-001" \
  -d "{
    \"transcript\": \"Hello, this is a format test.\",
    \"avatar_handle\": \"$SUME_AVATAR_HANDLE\",
    \"output_format\": {
      \"container\": \"wav\",
      \"sample_rate\": 24000,
      \"encoding\": \"pcm_s16le\"
    }
  }"

Limits that sit next to the format

A single Sume request takes up to 20,000 characters, and synthesized audio longer than 1200 seconds fails with tts_duration_exceeded without a credit capture. Check both against your script length before choosing a format, since a long WAV is large. A mono 24 kHz 16-bit WAV runs at 48,000 bytes per second, about 2.9 MB per minute, while the mp3 default at 128 kbps is 16,000 bytes per second, about 0.96 MB per minute. That is plain arithmetic, not a measured file size, so verify on a real clip. Current prices are listed by GET /v1/catalog, described in the API reference.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume