Phone-quality TTS: 8 kHz mu-law on Sume and where MAI Flash fits
Sume's TTS output_format accepts 8000 Hz and pcm_mulaw in wav or raw. What that gives a phone system, what it does not, and how MAI Flash is positioned.

Yes, Sume's TTS schema lets you ask for phone-style audio. output_format takes a container of mp3, wav or raw, a sample rate that includes 8000 Hz, and for wav and raw an encoding that includes pcm_mulaw and pcm_alaw. So a prompt for a telephony system can be requested as 8 kHz mu-law directly. The default is different: mp3 at 44,100 Hz and 128 kbps. Test one prompt on your own phone path before you render a library.
What the schema allows
These values are from the OpenAPI description of POST /v1/tts-1.0/generate. Not every combination is guaranteed to be offered by every voice, so verify by rendering one file.
| Option | Sume TTS 1.0 | OpenAI TTS guide |
|---|---|---|
| Containers or formats | mp3, wav, raw | MP3 (default), Opus, AAC, FLAC, WAV, PCM |
| Sample rates | 8000, 16000, 22050, 24000, 44100, 48000 | Not listed on the page read |
| Telephony encodings | pcm_mulaw, pcm_alaw (wav or raw) | Not listed on the page read |
| Streaming | Not streaming: async job with poll or webhook | Chunk transfer; WAV or PCM recommended for lowest delay |
How to request it
Send the encoding explicitly, and use wav if you also want sentence slices, because slice audio_url values are produced only for wav or raw.
{
"transcript": "Thanks for calling. Press one for sales.",
"avatar_handle": "@your_voice",
"language": "en",
"output_format": {
"container": "wav",
"sample_rate": 8000,
"encoding": "pcm_mulaw"
}
}Where the call-center pitch fits
Microsoft positions MAI-Voice-2.1-Flash for call centers, voice assistants and IVR systems, at about 45 ms of model inference and $15 per million characters (read 2026-10-04). That pitch is about speed when the text is created during the call. If your calls play fixed prompts, speed is a smaller factor than file format and loudness, and a job that renders once is enough, as in the fixed prompts walkthrough.
If the text is generated live, Sume's TTS is the wrong shape: it is an async job you poll or receive by webhook, not a stream (see Jobs and results).
Checks for phone audio
Narrowband audio hides detail, so numbers and names can blur. Listen on a real handset, spell out digit strings with spaces or hyphens in the transcript, and keep volume in range with generation_config.volume (0.5 to 2.0) rather than clipping. If you join prompts into one file later, timeline audio requires parts that share a channel layout, so keep one format throughout.
Sources
Related posts
More in Developers
- Pick seedance-2 or seedance-2.5 in code by clip length, with price
A small Python helper that sends clips up to 15 s to seedance-2 and longer ones to seedance-2.5, with the 720p price from Sume token math. Runs as written.
- Pick a subtitle translation model with a 20-cue pilot
Vendor benchmark scores do not tell you how a model handles your captions. Run the same 20 cues through each candidate, burn them, and compare on screen.
- Pin AI music model ids: music_v2_5, v6, lyria-3.5 and sume/music-auto
Which model id to pin for ElevenLabs Music, Suno and Google Lyria, and what Sume does with sume/music-auto. A config table and a rule for logging the engine.
- 1080x1350 sent as aspect_ratio: what Ideogram, Grok, Imagen get
Send pixels instead of a ratio and Sume snaps to the nearest native ratio. Tested table for 1080x1350, 1200x628 and 1500x500 across four model families.
Written by Sume