MAI-Voice 24 kHz 160 kbps MP3 vs Sume TTS 44.1 kHz 128 kbps: mixing

Microsoft's MAI-Voice example saves 24 kHz 160 kbps mono MP3; Sume TTS defaults to 44.1 kHz 128 kbps MP3 or WAV. What the numbers mean for mixing and file size.

5 min readSume
All posts

The Learn page's MAI-Voice examples request audio-24khz-160kbitrate-mono-mp3, a 24 kHz mono MP3 at 160 kbps; Sume's TTS 1.0 defaults to an MP3 at 44.1 kHz and 128 kbps and can return WAV. The higher MP3 bitrate does not make the Microsoft file better: 24 kHz sampling cannot carry content above 12 kHz, while 44.1 kHz reaches 22.05 kHz. For speech, both sound fine; for a video mix, the 44.1 kHz file avoids resampling to the usual video rates.

This is a format note, not a quality ranking of the two voices. I compared specifications, not audio.

The numbers

File size follows from bitrate. 160 kbps is 20,000 bytes a second, 1.2 MB a minute. 128 kbps is 16,000 bytes a second, 0.96 MB a minute. Uncompressed 16-bit mono WAV at 44.1 kHz is 88,200 bytes a second, about 5.3 MB a minute (stereo would double it; Sume's WAV channel count is not something I checked).

Microsoft's examples are one choice from a list of output formats; the Speech SDK can set others, so the 24 kHz MP3 is the documented sample, not a limit.

Output format comparison, docs examples (read 2026-10-08)
ItemMAI-Voice Learn exampleSume TTS 1.0 default
ContainerMP3MP3
Sample rate24 kHz44.1 kHz
Bitrate160 kbps128 kbps
ChannelsMonoNot stated in the contract excerpt
Per minute1.2 MB0.96 MB
AlternativeOther SDK output formatsWAV, pcm_s16le, 44.1 kHz

Choosing on Sume

Ask for WAV when the voice will be mixed with music and re-encoded once at the end; each extra MP3 generation adds artifacts. Ask for MP3 when the file is the deliverable, such as a podcast ad. Set it with output_format, for example {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100}.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"transcript": "Welcome back to the show.",
       "voice": {"mode": "id", "id": "YOUR_VOICE_ID"},
       "output_format": {"container": "wav", "encoding": "pcm_s16le",
                         "sample_rate": 44100},
       "mode": "async"}'

One caution

If a voice is later turned into speech-to-text input, a 16 kHz mono copy is enough and smaller. Keep the high-rate master for the mix and derive the small copy, rather than synthesizing twice.

Practical rule

Pick one working sample rate for a project and convert everything to it once, at the start. Video platforms commonly use 44.1 or 48 kHz for audio tracks, so a 24 kHz voice will be resampled at some point; doing it yourself with a good resampler beats leaving it to a platform's default. Keep the first-generation file untouched so you can redo the conversion.

Loudness matters more than sample rate for speech in a mix. Normalize the voice before you add music, and duck the music under it. A voice at the wrong level sounds worse than a voice at a slightly lower sample rate.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume