Phone-quality TTS: 8 kHz mu-law on Sume and where MAI Flash fits

Sume's TTS output_format accepts 8000 Hz and pcm_mulaw in wav or raw. What that gives a phone system, what it does not, and how MAI Flash is positioned.

5 min readSume
All posts

Yes, Sume's TTS schema lets you ask for phone-style audio. output_format takes a container of mp3, wav or raw, a sample rate that includes 8000 Hz, and for wav and raw an encoding that includes pcm_mulaw and pcm_alaw. So a prompt for a telephony system can be requested as 8 kHz mu-law directly. The default is different: mp3 at 44,100 Hz and 128 kbps. Test one prompt on your own phone path before you render a library.

What the schema allows

These values are from the OpenAPI description of POST /v1/tts-1.0/generate. Not every combination is guaranteed to be offered by every voice, so verify by rendering one file.

TTS output format options, Sume schema versus OpenAI guide (read 2026-10-04)
OptionSume TTS 1.0OpenAI TTS guide
Containers or formatsmp3, wav, rawMP3 (default), Opus, AAC, FLAC, WAV, PCM
Sample rates8000, 16000, 22050, 24000, 44100, 48000Not listed on the page read
Telephony encodingspcm_mulaw, pcm_alaw (wav or raw)Not listed on the page read
StreamingNot streaming: async job with poll or webhookChunk transfer; WAV or PCM recommended for lowest delay

How to request it

Send the encoding explicitly, and use wav if you also want sentence slices, because slice audio_url values are produced only for wav or raw.

{
  "transcript": "Thanks for calling. Press one for sales.",
  "avatar_handle": "@your_voice",
  "language": "en",
  "output_format": {
    "container": "wav",
    "sample_rate": 8000,
    "encoding": "pcm_mulaw"
  }
}

Where the call-center pitch fits

Microsoft positions MAI-Voice-2.1-Flash for call centers, voice assistants and IVR systems, at about 45 ms of model inference and $15 per million characters (read 2026-10-04). That pitch is about speed when the text is created during the call. If your calls play fixed prompts, speed is a smaller factor than file format and loudness, and a job that renders once is enough, as in the fixed prompts walkthrough.

If the text is generated live, Sume's TTS is the wrong shape: it is an async job you poll or receive by webhook, not a stream (see Jobs and results).

Checks for phone audio

Narrowband audio hides detail, so numbers and names can blur. Listen on a real handset, spell out digit strings with spaces or hyphens in the transcript, and keep volume in range with generation_config.volume (0.5 to 2.0) rather than clipping. If you join prompts into one file later, timeline audio requires parts that share a channel layout, so keep one format throughout.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume