MAI-Voice 24 kHz 160 kbps MP3 vs Sume TTS 44.1 kHz 128 kbps: mixing
Microsoft's MAI-Voice example saves 24 kHz 160 kbps mono MP3; Sume TTS defaults to 44.1 kHz 128 kbps MP3 or WAV. What the numbers mean for mixing and file size.

The Learn page's MAI-Voice examples request audio-24khz-160kbitrate-mono-mp3, a 24 kHz mono MP3 at 160 kbps; Sume's TTS 1.0 defaults to an MP3 at 44.1 kHz and 128 kbps and can return WAV. The higher MP3 bitrate does not make the Microsoft file better: 24 kHz sampling cannot carry content above 12 kHz, while 44.1 kHz reaches 22.05 kHz. For speech, both sound fine; for a video mix, the 44.1 kHz file avoids resampling to the usual video rates.
This is a format note, not a quality ranking of the two voices. I compared specifications, not audio.
The numbers
File size follows from bitrate. 160 kbps is 20,000 bytes a second, 1.2 MB a minute. 128 kbps is 16,000 bytes a second, 0.96 MB a minute. Uncompressed 16-bit mono WAV at 44.1 kHz is 88,200 bytes a second, about 5.3 MB a minute (stereo would double it; Sume's WAV channel count is not something I checked).
Microsoft's examples are one choice from a list of output formats; the Speech SDK can set others, so the 24 kHz MP3 is the documented sample, not a limit.
| Item | MAI-Voice Learn example | Sume TTS 1.0 default |
|---|---|---|
| Container | MP3 | MP3 |
| Sample rate | 24 kHz | 44.1 kHz |
| Bitrate | 160 kbps | 128 kbps |
| Channels | Mono | Not stated in the contract excerpt |
| Per minute | 1.2 MB | 0.96 MB |
| Alternative | Other SDK output formats | WAV, pcm_s16le, 44.1 kHz |
Choosing on Sume
Ask for WAV when the voice will be mixed with music and re-encoded once at the end; each extra MP3 generation adds artifacts. Ask for MP3 when the file is the deliverable, such as a podcast ad. Set it with output_format, for example {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100}.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"transcript": "Welcome back to the show.",
"voice": {"mode": "id", "id": "YOUR_VOICE_ID"},
"output_format": {"container": "wav", "encoding": "pcm_s16le",
"sample_rate": 44100},
"mode": "async"}'One caution
If a voice is later turned into speech-to-text input, a 16 kHz mono copy is enough and smaller. Keep the high-rate master for the mix and derive the small copy, rather than synthesizing twice.
Practical rule
Pick one working sample rate for a project and convert everything to it once, at the start. Video platforms commonly use 44.1 or 48 kHz for audio tracks, so a 24 kHz voice will be resampled at some point; doing it yourself with a good resampler beats leaving it to a platform's default. Keep the first-generation file untouched so you can redo the conversion.
Loudness matters more than sample rate for speech in a mix. Normalize the voice before you add music, and duck the music under it. A voice at the wrong level sounds worse than a voice at a slightly lower sample rate.
Sources
Related posts
More in Media tools
- Music from a still: image_url conditioning, still $0.125
The Music Router accepts an optional public HTTPS image_url to condition the track on a still. Request example, keeping a score consistent, fixed price.
- Remove backgrounds from 500 product photos: RMBG API cost and code
Sume RMBG 1.0 removes a background for $0.0225 per image regardless of size. Batch cost for 50 to 5,000 photos, the request body and a Python loop.
- TikTok trending search: 10 or 50 results, both $0.10
A Sume trending-videos search is $0.10 per accepted call whatever the limit. In production the limit runs 1 to 50 with a default of 10. Fields and a curl.
- Transcribe a video in one call: video-inspect with sentence segments
Skip detach: video-inspect with transcribe true returns a transcript with words and sentence segments, at $0.01 per audio minute plus the inspect's compute.
Written by Sume