Text to speech mulaw 8000 Hz: Gemini 3.8 and Sume TTS
Sume TTS can return pcm_mulaw or pcm_alaw at 8000 Hz in wav or raw containers. Gemini 3.8 TTS does it with audio/mulaw and audio/alaw mime types.

Sume TTS 1.0 can output mu-law or A-law at 8000 Hz: set output_format to a wav or raw container, encoding to pcm_mulaw or pcm_alaw, and sample_rate to 8000. The default is mp3, so you have to ask. Gemini 3.8 TTS offers the same codecs through audio/mulaw and audio/alaw mime types.
Gemini facts are from its speech-generation guide; Sume facts from the API reference. Read 2026-10-01.
How does Gemini ask for telephony audio?
The guide lists audio/mulaw (8-bit G.711 mu-law, common in North American and Japanese telephony) and audio/alaw (8-bit G.711 A-law, common in European and international telephony), set in response_format.mime_type, with an optional sample_rate such as 8000.
| Item | Gemini 3.8 TTS | Sume TTS 1.0 |
|---|---|---|
| Selector | mime_type: audio/mulaw, audio/alaw | output_format.encoding: pcm_mulaw, pcm_alaw |
| Sample rate | For example 8000 | One of 8000, 16000, 22050, 24000, 44100, 48000 |
| Container | Implied by mime type | wav or raw for PCM encodings |
| Default | WAV (unary) | mp3, 44100 Hz |
What does the Sume request look like?
Only the output_format object changes. Encoding applies to wav and raw containers, so an mp3 request cannot carry mu-law.
{
"output_format": {
"container": "wav",
"encoding": "pcm_mulaw",
"sample_rate": 8000
}
}What should I check before using it on a phone line?
Play the file through your carrier or PBX once; the reference does not describe telephony testing. Pick mu-law or A-law to match your trunk. For prompt content and structure, see IVR and voicemail prompts with text to speech.
Sources
Related posts
More in Developers
- Gemini TTS voice design voice_ id vs Sume voice ids
Gemini voice design returns a persistent voice_ id from a text prompt. Sume TTS accepts only a voice UUID or voi_ library id, and rejects other shapes with 400.
- GPT Image 2 input_fidelity: omit it; Sume returns 400
OpenAI says to omit input_fidelity for gpt-image-2 because inputs run at high fidelity. Sume lists no such field and rejects unlisted parameters with 400.
- GPT Image moderation_blocked vs Sume content_policy_rejected
OpenAI returns moderation_blocked with moderation_details. On Sume, policy refusals are grouped under content_policy_rejected. What to read and when to retry.
- GPT Image revised_prompt: what a Sume images response returns
OpenAI returns revised_prompt on the image generation call. A Sume POST /v1/images response has data[].url and usage, with no revised prompt field.
Written by Sume