Text to speech sample rate: which output_format to pick

Sume's TTS output_format takes mp3, wav or raw, six sample rates from 8000 to 48000 Hz and four PCM encodings. The default, and when to send wav.

4 min readSume
All posts

If you send no output_format, Sume's text to speech API returns an mp3 at 44100 Hz and 128 kbps. Send wav with pcm_s16le at 44100 Hz when the audio will feed an avatar mux or when you want sentence slices. The container can be mp3, wav or raw; the sample rate is one of 8000, 16000, 22050, 24000, 44100 or 48000 Hz.

All values are from the TTS request schema in the Sume API reference, read 2026-09-29, and apply to TTS 1.0 and the TTS Router alike.

What can output_format contain?

output_format is an optional object with four fields. Each is an enum in the schema.

From the TTS request schema in the Sume API reference, read 2026-09-29.
FieldAllowed valuesNote in the schema
containermp3, wav, rawAudio container
sample_rate8000, 16000, 22050, 24000, 44100, 48000Sample rate in Hz
bit_rate32000, 64000, 96000, 128000, 192000MP3 bit rate; required for mp3 when overriding defaults
encodingpcm_f32le, pcm_s16le, pcm_mulaw, pcm_alawPCM encoding for wav and raw containers

When should I send wav instead of mp3?

The schema names two cases. The first is feeding the audio into an avatar mux: pass wav, pcm_s16le and 44100 explicitly, because the default is mp3. The second is sentence segmentation. segmentation.mode=sentence needs timestamps.words: true, and each segment gets a sample-exact audio_url only when the container is wav or raw. With mp3 you still get the timings but no slice URLs.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tts-wav-001" \
  -d '{
    "transcript": "Your order ships tomorrow. Tracking follows by email.",
    "avatar_handle": "acme",
    "language": "en",
    "output_format": {
      "container": "wav",
      "sample_rate": 44100,
      "encoding": "pcm_s16le"
    },
    "timestamps": { "words": true },
    "segmentation": { "mode": "sentence" }
  }'

Which combination should I send?

Nothing in the schema ranks the sample rates, so choose by where the file goes. These rows only restate what the schema and its descriptions say.

From the TTS request schema in the Sume API reference, read 2026-09-29.
GoalSendWhy the schema says so
Just get audio backOmit output_formatDefaults to mp3 44100 Hz, 128 kbps
Feed an avatar muxwav, pcm_s16le, 44100The schema says to pass these explicitly for the avatar mux
Sentence audio sliceswav or raw, with timestamps.words and segmentation.mode=sentenceSegment audio_url slices come only from wav or raw
Pick an mp3 bit ratemp3 with a bit_rate from 32000 to 192000bit_rate is the mp3 setting

What about telephony sample rates?

The lowest rate on the list is 8000 Hz, and the encodings include pcm_mulaw and pcm_alaw. The schema does not say what any phone system expects, so check your provider's audio spec before you pick one; the docs promise only that these values are accepted.

What limits still apply?

  • The output_format schema sets additionalProperties: false, so send only the four fields above.
  • On TTS 1.0, synthesized audio longer than 1200 seconds fails with tts_duration_exceeded, whatever the format.
  • A bit_rate is for mp3. The schema describes encoding as the PCM setting for wav and raw.
  • For the endpoints and their prices, see Text to speech API.
  • For the word timings themselves, see Text to speech with highlighted words.

Sources

Related posts

More in Models

All Models posts

Written by Sume